AMAZINGINDEX.COM 日报快照
52.8
VOL. 2026.07
2026.07.31
← 返回 2026.07.31 日报
日报快照 · Daily Snapshot
NO. 016

GPT-5.6 代理创业 24 小时亏损 $99

#ARTICLE HackerNews 2026.07.31
推荐指数 39.0 NO. 016 · 2026.07.31
发布2026/07/30Score140Comments83

研究团队让 GPT-5.6 Sol 在真实环境中自主运营一家初创公司 24 小时,配备钱包、服务器和完整工具链。结果:消耗 3.2 亿 token、执行 1129 次工具调用,净亏损 $99.5,零收入,仅增 5 个用户。

这个实验的真正价值在于量化了当前 agent 的'幻觉成本'——908 次 shell 调用中大量是无效操作,说明长周期自主执行时的错误累积问题远比单次任务严重。团队公开了完整日志,这是目前少有的'真实业务场景'基准测试,比 GAIA、SWE-bench 更能反映落地难度。

对于正在做 agent 框架的创业者,这个数据点很关键:当前 frontier model 的可靠性天花板,可能意味着'完全无人值守'的商业模式在 2025 年还不成立,半自主+人工审核的混合架构仍是更务实的选择。

负面 77 条评论

核心争论:AI代理自主创业是噱头还是真具商业能力?实验设计是否严谨?

recitedropper

Pair this with the Hugging Face incident, and it hints that OpenAI is currently training their models to aggressively reward hack. That doesn't feel like a good sign to me--for the AI bull or the AI bear cases.

skybrian

They are being trained to try lots of unlikely alternatives and to be persistent. This often works well when searching for security bugs or counterexamples to famous math conjectures. But maybe it doesn't work so well when caution is required?

dylan604

"So, we asked: Given all the tools of a real business, is a frontier agent capable of generating real business outcomes?" "It Lied, Spammed, and Lost $447." Sounds like a vast majority of VC startups to me. From growth hacking to God views to all of the other disruption excuses, it just feels natura

替代方案: Hugging FaceComfyUI-MCPClaude vending machine experiment
查看原文 →