AMAZINGINDEX.COM 日报快照
58.6
VOL. 2026.07
2026.07.23
← 返回 2026.07.23 日报
日报快照 · Daily Snapshot
NO. 018

LLM 被怀疑刷榜"鹈鹕骑车"测试

#ARTICLE HackerNews 2026.07.23
推荐指数 53.0 NO. 018 · 2026.07.23
发布2026/07/22Score78Comments25

开发者用"画一只骑自行车的鹈鹕"这个经典 prompt 测试各大模型,发现结果异常完美,引发 AI 实验室可能针对性优化该 benchmark 的猜测。这暴露了非正式评测的脆弱性,做模型选型的工程师需要警惕。

这个梗能火恰恰说明行业评测体系的真空。正经的 LLM 评测要么太学术(MMLU、HumanEval 跟实际用感脱节),要么太容易被污染(训练数据里灌了测试题)。"鹈鹕骑车"这种野生 benchmark 反而成了用户信任的锚点,但一旦被针对性优化就彻底失效。

做模型选型的团队应该建立自己的私有评测集,用业务真实 prompt 做盲测,而不是追公开 benchmark 的分数。Anthropic 和 OpenAI 的模型在公开榜单上互有胜负,但在具体代码生成场景里差距可能完全相反。

意见分歧 27 条评论

核心争论:刷榜嫌疑是否成立,还是模型泛化能力提升的自然结果

dcchambers

It's incredible that each model has it's own style that remains relatively consistent throughout all of the different generated examples.

andy99

If an AI researcher was going to pelicanmaxx, they would almost certainly apply the augmentations mentioned in the article during training, e.g. randomly selecting animals and conveyances. You’d want a model that generalizes well, just sfting in that specific prompt would be pretty bush league for a

cute_boi

At this point, I think there are so many pelican images in the pretraining data that drawing a pelican no longer makes sense as a model evaluation task.

替代方案: Recraft V4pairwise comparison with ELOman sitting in chair at computer behind desk
查看原文 →