20万美元悬赏爬取Google图书全库
推荐指数 43.0 NO. 017 · 2026.07.05
发布2026/07/04Score120Comments33
为什么值得看
Anna's Archive发布20万美元赏金,寻求能大规模获取Google Books扫描全本的技术方案,而非仅搜索片段。对AI训练数据饥渴的创业者和研究者而言,这可能是解锁数千万册绝版书的关键通道。
编辑判断
这个悬赏的实质是破解Google Books的'片段搜索'架构——Google故意只暴露上下文片段防止全本下载,但扫描件确实存在于服务器上。历史上类似的攻防发生在Sci-Hub和Elsevier之间,核心思路通常是利用内部API端点或批量请求模式识别。
如果你是做数据获取的工程师,值得研究Google Books的inurl:books?id=参数结构和内部预览服务的请求签名,而不是死磕前端搜索。另一个隐蔽路径是Google Partner Program中出版社上传的完整PDF,权限边界可能存在配置漏洞。
对AI创业者更现实的信号是:训练数据的灰色市场正在明码标价,20万美元意味着他们认为全本图书数据的市场价值至少在百万美元以上。
社区反馈
意见分歧 32 条评论
核心争论:AI数据饥渴与版权壁垒的冲突,以及中国免费模型是否为泡沫症状
One of my hopes is that when the AI bubble bursts, some brave person will sneak out a copy of the last frontier model.
Not worried about that, you will only have to wait 3-6 months and get a Chinese model just as good.
Chinese companies giving away expensive models for free is a symptom of the AI bubble, too. It's not a law of nature that they'll always be able to scrounge up the money for yet another training run.