AI编程工具成本差17倍,选错亏大钱
推荐指数 69.0 NO. 011 · 2026.09.03
发布2026/09/02Score55Comments29
为什么值得看
Runta平台对比9款AI编程harness(Codex、Claude Code、Kimi Code等),发现同一模型在不同harness下单次通过成本差异高达17倍。Claude Code通过率高但成本达$18.34,OpenCode看似便宜实则失败率极高,缓存命中率与真实成本脱钩。
编辑判断
这个benchmark戳破了行业一个常见误区:把缓存命中率当成本指标。很多团队选harness时只看缓存率或通过率,但实际账单取决于失败重试的隐性消耗。OpenCode的例子很典型,15次通过的成本被失败任务摊薄后暴涨到$3.24,比表面数字难看得多。
对于在用Claude Code或Codex的团队,建议直接拿自己的任务集去Runta跑一遍真实成本,而不是依赖厂商公布的基准数据。特别是做AI编程基础设施的创业者,这里有个明确的产品机会:做harness层面的智能路由,根据任务类型动态切换后端,把成本压到当前最优解的1/3以下。
社区反馈
意见分歧 29 条评论
核心争论:harness与模型如何解耦评估:应测harness能力还是模型能力,成本-通过率权衡是否公平
Software is constrained when you write it. Agents have to be constrained while they run.
Codex is the best harness to me because of its GUI and subscription.
It's nice to see time reflected here. Deepseek is fuckin fast! Seems like the bigger the task, the faster it gets, which is kind of unfortunate because nobody is going to be doing 17 benchmark passes on a $50-100 task. I'm assuming the brief pause before it avalanches out 16kb of text at upwards of