AMAZINGINDEX.COM 日报快照
50.0
VOL. 2026.07
2026.07.15
← 返回 2026.07.15 日报
日报快照 · Daily Snapshot
NO. 013

27B模型首次塞进手机

#ARTICLE HackerNews 2026.07.15
推荐指数 73.0 NO. 013 · 2026.07.15
发布2026/07/14Score65Comments14

Bonsai 27B基于Qwen3.6 27B,通过1-bit/三值量化将模型从54GB压缩到可运行在手机端,支持多步推理、工具调用和视觉任务。端侧部署27B级别模型的门槛被彻底打破,本地AI Agent和隐私敏感场景迎来关键拐点。

1-bit/三值量化之前被质疑的主要是推理质量损失和训练稳定性,Bonsai系列之前用商用级模型证明了可行性,这次27B的发布意味着极端量化路线已经能覆盖复杂Agent场景。这对Apple Intelligence和Gemini Nano的端侧策略形成直接压力——它们目前依赖的还是7B以下小模型或云端协同。

对创业者来说,两个机会窗口打开:一是本地长上下文隐私场景(医疗、法律、企业内网)可以不再妥协模型能力;二是量化工具链本身可能成为新层,类似GGUF在4-bit时代的地位。需要关注的是,1-bit模型对定制微调的支持度如何,以及ARM NPU的实际推理延迟是否跟得上宣传。

意见分歧 15 条评论

核心争论:1-bit量化实际有效位数与性能损失的权衡,端侧部署是否值得

alvatech

TIL that 1 bit models are actually 1.58 bit with three values +1, 0 and -1

bensyverson

Yeah, it's an unfortunate convention from the very first "1 bit" model. But to be clear, Bonsai comes in both ternary and actual 1-bit variants.

NitpickLawyer

There's two variants of this (or, as the joke goes, for very big values of bit): Ternary Bonsai 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, giving a true 1.71 effective bits per weight. 1-bit Bonsai 27B uses binary {−1, +1} weights with the same group-wise scaling, giving 1.12

替代方案: Ornith 9BGemma 4-31BQwen 3.6 35BQwen 3.5 9B35B-A3B
查看原文 →