AMAZINGINDEX.COM 日报快照
59.8
VOL. 2026.08
2026.08.13
← 返回 2026.08.13 日报
日报快照 · Daily Snapshot
NO. 016

Grok 4.6 重返模型第一梯队

#ARTICLE HackerNews 2026.08.13
推荐指数 58.0 NO. 016 · 2026.08.13
发布2026/08/12Score166Comments132

SpaceXAI 的 Grok 4.6 在 Artificial Analysis Intelligence Index 上得 61 分,追平 GPT-5.6 Sol,agentic 能力 Elo 1753 仅次于 Claude Opus 5。对 AI 工程师意味着多了一家可替代 OpenAI 的前沿模型供应商,且推理成本更低。

Grok 4.6 的迭代速度值得注意:一个月提升 5 分,三个月提升 23 分,这种追赶节奏和 SpaceX 的火箭迭代文化一脉相承。更关键的是它在 agentic 场景以更低成本逼近 Claude Opus 5,这对正在做 AI Agent infra 的团队是实质性利好——你可以用 Grok API 替代部分 Claude 调用,把预算腾给评估和编排层。

但别急着全量迁移。Grok 的历史版本在复杂推理一致性上波动较大,建议先用 GDPval-AA 覆盖的场景跑一轮回归测试,特别关注长链工具调用时的错误累积率。如果验证通过,Grok 4.6 可能是目前性价比最高的 frontier agentic 模型。

意见分歧 106 条评论

核心争论:Grok 4.6 性价比是否足以替代 OpenAI/Anthropic,还是智力差距仍关键

satvikpendem

Cursor, since Grok 4.5, has had an incredible deal for frontier level models, their subscription now goes way further than OpenAI or Anthropic. Even on their lower tier plans you can use a lot tokens on their of their first party models (Grok and Composer) and not really run out comparatively. Combi

aliljet

Can you explain what you mean? These days courtesy of an addictive reset game OpenAI is playing, I can't find anything with frontier intelligence that's more cost efficient...

timr

If they didn’t constantly reset, they’d be about the same as Anthropic. Right now, I find that Grok offers better value, uses fewer tokens per turn, and makes better code. I haven’t tried Cursor because I don’t want to change editors again, but maybe I should try it…

替代方案: OpenAIAnthropicClaude Opus 5GPT-5.6 SolCursorGitHub CopilotKimi K3Fable 5Opencodetyped++GabAIMuse Spark Contributor Tier
查看原文 →