Stability AI's family of licensed-data audio models for generating instrumental music and sound effects from text — now the center of the company's music-first pivot.
Last verified · 2026-10-04 · by Moe Ameen
Stable Audio is Stability AI's family of generative audio models that turn a text prompt into instrumental music or sound effects. It is the product Stability is betting its future on: reporting in early October 2026 (Sean Parker to The Information, covered by TechCrunch on October 2, 2026) describes Napster co-founder and Stability backer Sean Parker and CEO Prem Akkaraju rebuilding the maker of Stable Diffusion into an AI toolmaker for professional musicians, anchored by a $76 million round from Universal, Sony, Warner, and EA in which the three major labels also licensed their catalogs to Stability for training.
The product predates the pivot. Stable Audio 2.5, launched in September 2025, was pitched as the first audio model built for enterprise sound production — it generates tracks up to about three minutes in under two seconds on a GPU, supports audio inpainting (hand it a clip, pick a point, and it continues the track), and produces structured compositions with an intro, development, and outro. In May 2026, Stability released Stable Audio 3.0, a four-model family (Small SFX, Small, Medium, Large) that generates tracks up to about six minutes, was trained entirely on licensed data, and adds multi-segment inpainting and LoRA fine-tuning.
What sets it apart from other generators is twofold. First, three of the four Stable Audio 3.0 models ship with open weights under the Stability AI Community License, so developers can download, run, and fine-tune them — the small models are a few hundred million parameters and produce clips in well under a second on a high-end GPU. Second, it is trained on licensed data, which is the piece that makes the "safe to use" pitch credible and is exactly what the label investment is meant to deepen.
One honest scope note: Stable Audio makes instrumental music and sound effects, not sung songs with lyrics. If you want a finished vocal track from a prompt, that is Suno or Udio. Stable Audio is built to produce the score, the bed, the hook, and the SFX — the audio you build something else on top of.
Think of Stable Audio as the scoring stage and [Kompozy](/) as the stage that wraps a finished, published video around that score. Stable Audio's real win for a creator is the soundtrack problem: a fast, licensed-data way to get a hook or a bed that won't trip copyright claims on Reels, TikTok, or YouTube. But a track on its own reaches no one — audio is an ingredient, and the video has to exist, be captioned, be sized per platform, and actually get posted. That second half is the job Kompozy automates.
The concrete workflow: generate your instrumental or hook in Stable Audio, then bring your source material into Kompozy and let it build the video that rides on top. Kompozy generates [Clipped Shorts](/glossary/clipped-short) from long-form footage, [Persona Shorts](/glossary/persona-shorts) and HeyGen avatar video from a script, [Listicle Video](/glossary/output-buckets) and Marketing Shorts — and your Stable Audio track becomes the audio bed under them, with word-synced captions burned in for muted feeds and the frame cut to 9:16, 1:1, and 16:9. A [Persona Brief](/glossary/persona-brief) keeps one voice across the batch, and [autopilot](/glossary/autopilot) schedules and publishes across the eight social platforms plus blog and email through a per-post review step. Kompozy runs its own stack (Claude and OpenAI for copy, HeyGen for avatars, gpt-image and Gemini for images) and does not generate music — Stable Audio owns the sound, Kompozy owns the finished, scheduled video around it.
Stable Audio is Stability AI's family of generative audio models that turn text prompts into instrumental music and sound effects. It spans an enterprise model (Stable Audio 2.5, September 2025) and an open-weight family (Stable Audio 3.0, May 2026, with three of four models downloadable), all trained on licensed data. It is accessed through the StableAudio.com web app, a credit-based API, and partner platforms, and is now the center of Stability's music-first repositioning under Sean Parker and CEO Prem Akkaraju.
No. Stable Audio is built for instrumental music and sound effects — full instrumental tracks or shorter snippets from a prompt — not sung songs with lyrics. For a finished vocal song from a text prompt, Suno or Udio are the right tools. A planned Stable Audio update will reportedly let you steer generation by humming a melody or beatboxing a drum pattern, but that guides the music, not lyric vocals, and had no release date at the time of writing.
Partly. The StableAudio.com web app is freemium (free to start, paid tiers above that), and three of the four Stable Audio 3.0 models — Small SFX, Small, and Medium — are free to download under the Stability AI Community License, which permits commercial use subject to registration and a revenue threshold. The developer API is paid per generation in credits (roughly $0.20 for a 2.5 track, about $0.26 for a 3.0 Large track). Confirm current figures at platform.stability.ai.
Stable Audio generates the audio but publishes nothing. Generate a cleared instrumental or hook, then bring your footage or script into Kompozy, which builds the video around it — Clipped Shorts, persona/avatar video, Listicle or Marketing Shorts — uses your track as the audio bed, burns in captions, reframes for each platform, and schedules and publishes across the eight social platforms plus blog and email from one queue.