Stability AI · 2026-05-20 · major
Stable Audio 3.0 — Stability AI Ships a Four-Model Audio Family With 6:20 Music Generation and Open Weights for Three Sizes
Stability AI launches Stable Audio 3.0, a family of four music and SFX models. Medium and Large generate full 6:20 compositions; Small SFX, Small, and Medium ship as open weights on Hugging Face.

A four-model audio family from Stability AI with three open-weight checkpoints and full 6:20 song generation in the larger sizes.
Key specs
| Max length | 6:20 |
|---|---|
| Model count | 4 |
| Open weight models | 3 |
| Small sfx params | 459M |
| Small params | 459M |
| Medium params | 1.4B |
| Large params | 2.7B |
| Training tracks | 806,284 |
What is it?
Stable Audio 3.0 is Stability AI's third-generation text-to-audio family. It includes Small SFX (sound effects), Small (music up to 2 minutes, on-device), Medium (music up to 6:20), and Large (highest quality, API only). It replaces 2024's Stable Audio 2.0 and more than doubles the maximum song length over the prior open Stable Audio Open model.
How does it work?
The models use a novel semantic-acoustic autoencoder paired with a text conditioner. They support variable-length generation, audio inpainting (single-segment, multi-segment, and causal continuation editing), and LoRA fine-tuning on user libraries. Training used 806,284 tracks from AudioSparx plus filtered Creative Commons material, all fully licensed.
Why does it matter?
Open-weight music generation was previously capped at clip-length outputs (10–47 seconds). Shipping a 6:20 open Medium model means hobbyists and indie tools can render full songs locally without API fees, and the Small models bring music and SFX generation to mobile and consumer laptops.
Who is it for?
musicians, indie developers, audio researchers, mobile-app makers
Try it
huggingface.co/collections/stabilityai/stable-audio-3