AMAZINGINDEX.COM 日报快照
55.2
VOL. 2026.07
2026.07.09
← 返回 2026.07.09 日报
日报快照 · Daily Snapshot
NO. 020

Cognition 新模型挑战后训练天花板

#ARTICLE HackerNews 2026.07.09
推荐指数 45.0 NO. 020 · 2026.07.09
发布2026/07/08Score137Comments86

Cognition 发布 SWE-1.7,基于 Kimi K2.7 继续 RL 训练后达到 GPT-5.5/Opus 级别智能,成本大幅降低。关键信号:在已有大量 RL 的基座上仍能大幅提升,说明 RL scaling 远未触顶,对做 agentic 工程的公司有直接参考意义。

Cognition 选择 Kimi K2.7 而非自研基座做 RL 续训,说明国内大模型的基座能力已被海外顶级 lab 认可为可投入生产的起点。这比模型本身更值得注意——Moonshot 的 toB 出海和 API 生态可能因此获得意外杠杆。

对读者的直接启示:如果你在做垂直 agent,不必执念自研基座,找已有强 RL 基础的模型做 domain-specific 续训,可能是更务实的路径。Cognition 的 infra 和数据 trick 细节(长程任务、稳定性)预计会在未来几个月逐步开源或论文化,建议跟踪其技术博客而非只看产品发布。

负面 72 条评论

核心争论:Cognition 声称的模型性能突破是否可信,还是典型的 VC 导向炒作

harmonic18374

A company whose first demo was completely fraudulent announces that its model beats GPT-5.5, on its own benchmark? I’m gonna wait a little before I trust this. This whole company seems to optimize for raising money and impressing VCs. Lying about their products, ignoring consumer market to target en

giancarlostoro

Would love to see these companies use benchmarks done by third parties.

anthonypasq

they are right there? it shows swe-bench multilingual and terminal bench

替代方案: GPT-5.5OpusKimi K2.7AnthropicOpenAI
查看原文 →