340+地方新闻站屏蔽互联网档案馆
推荐指数 44.0 NO. 020 · 2026.05.22
发布2026/05/21Score105Comments27
为什么值得看
McClatchy、Advance Local 等连锁报业集团加入屏蔽 Internet Archive 的行列,此前纽约时报等已因担忧 AI 公司爬取存档内容而限制访问。这对依赖 Archive 做数据回溯的 AI 训练团队和研究者构成直接障碍,公开网络语料的"免费午餐"时代正在加速终结。
编辑判断
这波屏蔽潮的真正目标不是 Archive 本身,而是借道 Archive 规避 robots.txt 限制的 AI 爬虫。OpenAI、Anthropic 等被曝多次从 Archive 的 Wayback Machine 和 Open Library 批量提取内容,因为 Archive 的服务器 IP 通常不在出版商的黑名单上。
对做预训练数据清洗的团队来说,这意味着两条路在收窄:一是通过第三方镜像获取历史内容的灰色渠道正在合法化收缩,二是出版商正在把"反 AI 爬取"从 robots.txt 升级到法律诉讼层面。纽约时报诉 OpenAI 的判例如果落地,会直接改变数据许可的商业模式。
建议现在就开始评估训练数据中的 Archive 来源占比,同时关注 DMCA 1201 条款在这类场景中的司法解释动向。已经在用 Archive 数据做 RAG 或微调的项目,需要准备替代数据源或正式的授权谈判。
社区反馈
意见分歧 28 条评论
核心争论:新闻站屏蔽互联网档案馆:保护版权合理还是加速信息私有化和历史抹除
Maybe they should allow the Internet Archive access to their article after a week or 2. But I think this will hurt them as time goes on more then help. IIRC, one news org blocked free access and their revenue fell. I think that was in Australia. But seems they are using AI as the reason. So allow
That sounds like a good idea to me. One of the tests for Fair Use in the US, as I understand it, would be whether the archived work "competes" with the original. If people start going to IA instead to read the news, the newspaper might have a claim. But if they're doing it to get around paywalls, or
In general judges seem to understand that the copyright holder has some interest in these situations but not seem to understand that the rest of the community has some rights too.