// AI TOOLS · STABLE AUDIO

Stable Audio

Stability AI's family of licensed-data audio models for generating instrumental music and sound effects from text — now the center of the company's music-first pivot.

Last verified · 2026-10-04 · by Moe Ameen

What Stable Audio is

Stable Audio is Stability AI's family of generative audio models that turn a text prompt into instrumental music or sound effects. It is the product Stability is betting its future on: reporting in early October 2026 (Sean Parker to The Information, covered by TechCrunch on October 2, 2026) describes Napster co-founder and Stability backer Sean Parker and CEO Prem Akkaraju rebuilding the maker of Stable Diffusion into an AI toolmaker for professional musicians, anchored by a $76 million round from Universal, Sony, Warner, and EA in which the three major labels also licensed their catalogs to Stability for training.

The product predates the pivot. Stable Audio 2.5, launched in September 2025, was pitched as the first audio model built for enterprise sound production — it generates tracks up to about three minutes in under two seconds on a GPU, supports audio inpainting (hand it a clip, pick a point, and it continues the track), and produces structured compositions with an intro, development, and outro. In May 2026, Stability released Stable Audio 3.0, a four-model family (Small SFX, Small, Medium, Large) that generates tracks up to about six minutes, was trained entirely on licensed data, and adds multi-segment inpainting and LoRA fine-tuning.

What sets it apart from other generators is twofold. First, three of the four Stable Audio 3.0 models ship with open weights under the Stability AI Community License, so developers can download, run, and fine-tune them — the small models are a few hundred million parameters and produce clips in well under a second on a high-end GPU. Second, it is trained on licensed data, which is the piece that makes the "safe to use" pitch credible and is exactly what the label investment is meant to deepen.

One honest scope note: Stable Audio makes instrumental music and sound effects, not sung songs with lyrics. If you want a finished vocal track from a prompt, that is Suno or Udio. Stable Audio is built to produce the score, the bed, the hook, and the SFX — the audio you build something else on top of.

What you can make with it

  • Cleared instrumental beds and background tracks for video, podcasts, and ads, trained on licensed data
  • Short musical hooks and stingers for the first few seconds of a short, where retention is decided
  • Sound effects, risers, and loops via the dedicated Small SFX model
  • Structured multi-part instrumentals (intro, development, outro) up to several minutes long
  • Extended or edited tracks using audio inpainting to continue or refine a specific segment
  • Self-hosted, fine-tuned audio from the open-weight 3.0 models, adapted to your own sound library via LoRA

How Kompozy turns Stable Audio output into content

Think of Stable Audio as the scoring stage and [Kompozy](/) as the stage that wraps a finished, published video around that score. Stable Audio's real win for a creator is the soundtrack problem: a fast, licensed-data way to get a hook or a bed that won't trip copyright claims on Reels, TikTok, or YouTube. But a track on its own reaches no one — audio is an ingredient, and the video has to exist, be captioned, be sized per platform, and actually get posted. That second half is the job Kompozy automates.

The concrete workflow: generate your instrumental or hook in Stable Audio, then bring your source material into Kompozy and let it build the video that rides on top. Kompozy generates [Clipped Shorts](/glossary/clipped-short) from long-form footage, [Persona Shorts](/glossary/persona-shorts) and HeyGen avatar video from a script, [Listicle Video](/glossary/output-buckets) and Marketing Shorts — and your Stable Audio track becomes the audio bed under them, with word-synced captions burned in for muted feeds and the frame cut to 9:16, 1:1, and 16:9. A [Persona Brief](/glossary/persona-brief) keeps one voice across the batch, and [autopilot](/glossary/autopilot) schedules and publishes across the eight social platforms plus blog and email through a per-post review step. Kompozy runs its own stack (Claude and OpenAI for copy, HeyGen for avatars, gpt-image and Gemini for images) and does not generate music — Stable Audio owns the sound, Kompozy owns the finished, scheduled video around it.

  1. Generate a cleared instrumental bed or a short hook in Stable Audio from a text prompt, using inpainting to extend or tighten it to the length you need.
  2. Bring your footage or script into Kompozy as a source and set a Persona Brief so every output holds one voice and your banned words.
  3. Generate the video: Clipped Shorts from long-form footage, a Persona Shorts or HeyGen avatar clip from a script, or Listicle/Marketing Shorts — auto-captioned and reframed to 9:16, 1:1, and 16:9.
  4. Drop your Stable Audio track in as the audio bed, and spin the strongest moments into a brand-exact Carousel and Quote Graphics, plus a Blog Article and Email Newsletter from the transcript.
  5. Schedule and publish the whole batch across the eight social platforms plus blog and email from one queue with autopilot and a per-post review pass.

Frequently asked questions

What is Stable Audio?

Stable Audio is Stability AI's family of generative audio models that turn text prompts into instrumental music and sound effects. It spans an enterprise model (Stable Audio 2.5, September 2025) and an open-weight family (Stable Audio 3.0, May 2026, with three of four models downloadable), all trained on licensed data. It is accessed through the StableAudio.com web app, a credit-based API, and partner platforms, and is now the center of Stability's music-first repositioning under Sean Parker and CEO Prem Akkaraju.

Can Stable Audio make a full song with vocals?

No. Stable Audio is built for instrumental music and sound effects — full instrumental tracks or shorter snippets from a prompt — not sung songs with lyrics. For a finished vocal song from a text prompt, Suno or Udio are the right tools. A planned Stable Audio update will reportedly let you steer generation by humming a melody or beatboxing a drum pattern, but that guides the music, not lyric vocals, and had no release date at the time of writing.

Is Stable Audio free?

Partly. The StableAudio.com web app is freemium (free to start, paid tiers above that), and three of the four Stable Audio 3.0 models — Small SFX, Small, and Medium — are free to download under the Stability AI Community License, which permits commercial use subject to registration and a revenue threshold. The developer API is paid per generation in credits (roughly $0.20 for a 2.5 track, about $0.26 for a 3.0 Large track). Confirm current figures at platform.stability.ai.

How do I turn a Stable Audio track into social content?

Stable Audio generates the audio but publishes nothing. Generate a cleared instrumental or hook, then bring your footage or script into Kompozy, which builds the video around it — Clipped Shorts, persona/avatar video, Listicle or Marketing Shorts — uses your track as the audio bed, burns in captions, reframes for each platform, and schedules and publishes across the eight social platforms plus blog and email from one queue.

Related tools

  • Suno — The consumer AI music generator that writes a full song — lyrics, vocals, instrumentation, and mix — from a text prompt, now under a copyright cloud over how it was trained.
  • ElevenLabs — ElevenLabs is the leading AI voice platform — ultra-realistic text-to-speech, voice cloning, dubbing, sound effects, speech-to-text, and conversational voice agents.
  • Stable Diffusion — Stability AI's open family of text-to-image models — self-hostable, deeply customizable, and freshly backed by a $76M round from Universal, Sony, Warner, and EA.
  • Fish Audio — An AI voice platform for expressive real-time text-to-speech and fast voice cloning — with an open-source model family (Fish Speech) and a hosted flagship, S2.1 Pro, aimed at creators, developers, and enterprises.
  • Audionaut — A free, open-source multitrack audio editor and recorder for macOS, Windows, and Linux — built in C++ on JUCE, with local stem separation and AI-agent control over MCP.

← All AI tools · Get started →