Stable Audio 3:可变长度音频生成
推荐指数 46.0 NO. 013 · 2026.05.21
发布2026/05/20Score65Comments13
为什么值得看
Stability AI 发布第三代音频生成模型,支持按实际时长生成而非固定长度输出,并内置音频编辑的 inpainting 功能。对需要批量生成音效、BGM 的开发者,可显著降低推理成本和后期剪辑工作量。
编辑判断
音频生成领域长期被固定时长输出困扰——生成 30 秒音效也得跑完整首 3 分钟模型的计算量。Stable Audio 3 的 variable-length 设计直接按需求长度采样,推理成本与输出时长线性挂钩而非固定上限。
此前 ElevenLabs 的 Sound Effects 和 Meta 的 AudioGen 都不支持原生可变长度,用户只能靠后期裁剪或重复拼接。SA3 的语义-声学联合自编码器(semantic-acoustic autoencoder)把文本语义和波形声学压到同一隐空间,inpainting 时可以直接用文本指定替换哪段内容,比 Adobe 的 Project Music GenAI Control 的纯波形编辑更精确。
已开源 small/medium/large 三档模型,显存需求从 8GB 到 24GB 不等。做游戏音效、播客后期、短视频 BGM 的团队可以优先试 medium 档,生成 10 秒音效的延迟能压到 2 秒内。
社区反馈
意见分歧 14 条评论
核心争论:Stability AI 能否靠开放权重模式持续生存,还是已因人才流失和商业化失误走向衰落
Stability.ai is still around? I thought they died because they gave away everything for free with no revenue model. Emad trained a lot of really great models, but he just gave them away. This cost enormous sums of money. I wish for a world where Stability gave away the weights, but had a monetizatio
I wonder if we just discovered that we’re living in that world. :) (To fill in some gaps: they’ve consistently had a revenue model, first subscriptions to use their models commercially then fixed-floor cost per generation with revenue sharing with them)
There are plenty of companies with open weights.