本地跑LLM比API更贵
推荐指数 60.0 NO. 014 · 2026.05.18
发布2026/05/17Score234Comments198
为什么值得看
作者实测M5 MacBook Pro运行离线LLM的完整成本,发现设备折旧加电费后,每百万token成本高于OpenRouter等API服务。这对"本地更省钱"的普遍假设提出了直接挑战。
媒体预览
编辑判断
这个结论很多人直觉上难以接受,但作者的计算框架是对的:大多数人算本地成本只算电费,忽略了设备折旧和利用率。M5 Pro满负载跑推理时,芯片寿命加速损耗是隐性大头。
更关键的变量是利用率。如果你每天只跑几小时,摊销到每百万token的硬件成本会飙升;只有7x24高负载运行,本地才可能打平API。这解释了为什么云厂商的推理服务能持续降价——他们的GPU利用率是你的10倍以上。
正在考虑本地部署的团队,建议先用这个模型算清自己的实际利用率,再决定买设备还是买API。
社区反馈
意见分歧 160 条评论
核心争论:本地LLM成本计算是否应包含整机折旧,还是仅算增量成本
How much does your data privacy cost?
As stated in the analysis, thousands of dollars. That said, the smart thing to do is target smaller models (few billion parameters) and then use larger models for non-privacy tasks.
Will this cost structure always be this way and are there other benefits to not running your LLM on the cloud? E.g. Privacy Uptime Future cost structure controls This is a field that has moved very quickly. And it has moved in a direction to try to trap users into certain habits. But these habits mi