AMAZINGINDEX.COM 日报快照
51.8
VOL. 2026.07
2026.07.26
← 返回 2026.07.26 日报
日报快照 · Daily Snapshot
NO. 009

Kimi K3 网络攻击能力获英美联合评估

#ARTICLE HackerNews 2026.07.26
推荐指数 74.0 NO. 009 · 2026.07.26
发布2026/07/25Score117Comments36

英美 AI 安全机构联合测评 Moonshot AI 的 Kimi K3,在漏洞利用基准 ExploitBench 上表现突出,该模型将于 7 月 27 日开源权重。这是首次有开源权重模型在公开网络能力评估中被标记为高风险,安全团队需提前准备防御方案。

开源权重的高网络能力模型是安全界的噩梦场景:攻击者可以本地微调、去掉安全限制、针对特定漏洞批量生成利用代码。此前 GPT-4 级别的网络能力被封闭在 API 后面,至少还有访问日志和速率限制作为缓冲。

做安全防御的团队现在就要行动:更新威胁模型,假设攻击者拥有端到端自动化漏洞利用能力;做红队的则需要重新评估自己的价值——如果模型能自动生成 exploit,人工红队的效率优势还剩多少。Moonshot 选择开源的底气可能来自中美 AI 竞赛的压力,但代价是全球攻击面的一次结构性扩大。

意见分歧 34 条评论

核心争论:开源模型K3是否真接近闭源模型,还是评估方法低估了其实际能力

avaer

What's "Top U.S. Models"? Even if we know what the set is, it's not clear what the numbers actually mean. Is it min, max, average, weighted, median of the models? Prerelease or public, with or without safeguards? A bit frustrating to have this be hand waved in a report, the graphs might as well just

guessmyname

NIST has named the top U.S. models in their (full) report: https://www.nist.gov/system/files/documents/2026/07/17/CAISI... Spoiler: it’s OpenAI’s GPT-5.5 and Anthropic’s Mythos Preview [**]. [**] Reminder that Mythos Preview is a very different beast from

pama

What is the rationale for gpt-5.5 when gpt-5.6-sol exists?

替代方案: GPT-5.5GPT-5.6-solMythos PreviewGLM-5.2Opus 5Fable5Mythos5Opus5Opus 4.5
查看原文 →