// OPEN-SOURCE AI VIDEO GENERATION REVIEW

Kandinsky 6.0 Video review (2026): is the free open model with sound worth it?

An honest review of Kandinsky 6.0 Video, Sber's free MIT-licensed AI video model with synchronized audio — scores, pros and cons, and who it's actually for.

Last verified · 2026-10-10 · by Moe Ameen
The verdict
4.0 / 5

As a free, MIT-licensed, open-weight model that generates 5-second clips with synchronized audio, Kandinsky 6.0 Video is a genuinely notable release — audio-native generation you can run yourself and use commercially is rare. But it is a model, not a product: the output is short, modest-resolution, and silent-optional raw footage, and getting published content out of it means your own GPU or ComfyUI, plus captions, reframing, and a schedule. Rate it high as a building block and low as a finished workflow.

On October 6, 2026, Sber's Kandinsky Lab open-sourced Kandinsky 6.0 Video — a family of video models that generate clips with synchronized sound, released with code and weights under an MIT license. There are two lines, a 29-billion-parameter Pro and a 3-billion-parameter Lite, and the same generate-with-sound capability is also live for free inside Sber's GigaChat assistant. For a category where most open models output silent video, shipping lip-synced speech, ambience, and music in the same pass is the headline.

I run a competing content engine, so weigh that, but this review is about the model on its own terms, not a sales pitch. The honest question for a creator is not "is the research impressive?" — it clearly is. It is "what do I actually get, and what do I still have to do myself?" Kandinsky 6.0 Video gives you a free, commercially usable, audio-equipped clip generator. What it does not give you is anything resembling a finished post: the clips are about five seconds, the base resolution is modest (a separate super-resolution model upscales to Full HD), and running the open weights assumes a capable GPU or a ComfyUI setup.

So this page scores two different things that searchers conflate: the model as a generation primitive, where it is strong, and the model as a way to produce publishable content, where it is only a first step. Everything below reflects the release as documented on 2026-10-10; where quality claims against closed models are early or self-reported, I say so.

What Kandinsky 6.0 Video is

Kandinsky 6.0 Video is an open-weight text-to-video and image-to-video model family from Kandinsky Lab, the generative-media team at Sber. It generates roughly 5-second clips at 24 fps and produces synchronized 44 kHz audio alongside the picture — covering speech with lip-sync, ambient effects, and music — with the option to generate silent video instead. It ships as a 29B Pro line and a 3B Lite line, each in pretrained and distilled (faster, fewer-step) variants, and a separate Kandinsky 6.0 super-resolution model upscales the modest base resolution to Full HD (1920×1080). Access is unusually open for a model with these capabilities. Weights are on Hugging Face under the `kandinskylab` organization, code is on GitHub, and the release arrived with Diffusers pipelines, a ComfyUI extension, vLLM-Omni support, a demo Space, and Colab/Kaggle notebooks. Because the license is MIT, the weights can be used commercially and self-hosted — you pay compute, not a per-clip fee. For non-technical users, the same capability is available free inside GigaChat, where the clip-length cap is currently about five seconds.

Who Kandinsky 6.0 Video is for

Kandinsky 6.0 Video fits technically comfortable creators and small teams who want cheap, repeatable, sound-equipped short clips and are happy to run a model locally or in ComfyUI — think B-roll, hooks, loops, and experimentation at volume without per-render metering. It is a weak fit for anyone who wants a finished, captioned, correctly-framed, scheduled post out of a single tool, or who needs clips longer than a few seconds, cinematic resolution out of the box, or a no-code path beyond GigaChat's free tier. If you are not going to touch a GPU or a node graph, the open weights are not really for you; the GigaChat path is.

Scoring breakdown

DimensionScoreWhy
Generation quality4.0 / 5Strong for an open 5-second model; Sber's own report says human reviewers preferred Pro over its predecessor, though head-to-head claims against closed models are early.
Synchronized audio4.5 / 5The standout: lip-synced speech, ambience, and music generated with the video in one pass, at 44 kHz — rare in open models.
Clip length & resolution3.0 / 5About 5 seconds at a modest base resolution; Full HD needs the separate upscaler, and longer clips are not yet supported.
Openness & licensing5.0 / 5MIT license on both code and weights, commercial use allowed, self-hostable — as open as it gets for a model this capable.
Ecosystem & integrations4.5 / 5Launched with Diffusers, ComfyUI, vLLM-Omni, a demo Space, and Colab/Kaggle notebooks — unusually complete at release.
Ease of use2.5 / 5Running the open weights assumes a capable GPU or ComfyUI. The free GigaChat path is easy but limited; there is no polished standalone app.
Value4.5 / 5Free weights and a free GigaChat tier make the cost compute-only; for high-volume short clips that is excellent value.
Publishing readiness2.0 / 5Output is raw footage — no captions, no reframing, no scheduling. Getting to a published post is entirely on you.

Pros and cons

Pros

  • Generates synchronized 44 kHz audio — lip-synced speech, ambience, and music — in the same pass as the video, which most open models cannot do.
  • MIT license on code and weights means genuine commercial use and self-hosting, with no per-clip fee beyond compute.
  • A 3B Lite line makes it runnable on a single high-end consumer GPU, not just data-center hardware.
  • Arrived with a real ecosystem — Diffusers, ComfyUI, vLLM-Omni, demo Space, Colab/Kaggle — so you can start the day it dropped.
  • A free GigaChat path lets non-technical users try generate-with-sound at no cost.
  • Silent-video option and image-to-video mode add flexibility beyond plain text-to-video.

Cons

  • Clips are only about five seconds — a single ingredient, not a finished post or a long-form video.
  • Base resolution is modest; Full HD requires running the separate super-resolution model.
  • Self-hosting the open weights needs a capable GPU and comfort with Diffusers or ComfyUI — not a no-code experience.
  • No built-in captioning, aspect-ratio reframing, scheduling, or publishing — all of that is on you.
  • Quality comparisons against closed models like Veo or Kling are early and self-reported.
  • The GigaChat free path is convenient but capped (about five seconds) and tied to Sber's assistant rather than a dedicated creator tool.

Pricing analysis

There is no price in the usual sense, which is the whole point. Kandinsky 6.0 Video's code and weights are MIT-licensed and free to download, so the only cost of running them is compute — a GPU you own or rent. For a creator generating a high volume of short clips, that is a materially different economic model than a hosted generator that meters per render: once you have the hardware, the marginal cost of another 5-second clip approaches zero. The 3B Lite line lowers the hardware bar enough that a single high-end consumer GPU is plausible, which is what makes the "free" claim real rather than theoretical.

The catch is that free weights are not free workflow. The time and skill to stand up Diffusers or ComfyUI, run the separate upscaler for Full HD, and then caption, reframe, and schedule the output are real costs that a per-clip hosted price would otherwise absorb. For non-technical creators, the free GigaChat tier removes the setup cost but adds limits — roughly five-second clips inside Sber's assistant rather than a dedicated creator surface.

Net: as a generation primitive, the value is excellent, especially at volume. As a path to finished, published content, "free" understates the real cost, because the model deliberately stops at the clip. Budget for the production and distribution layer separately — whether that is your own editing time or a tool that does it.

Use-case fit

Use caseFitWhy
High-volume short B-roll and hooksStrongFree, self-hostable, audio-equipped 5-second clips are ideal raw material to generate at volume.
Adding synchronized sound to AI videoStrongLip-synced speech, ambience, and music in one pass is exactly what this model is built for.
Commercial use on a budgetStrongThe MIT license permits commercial use with no per-clip fee; you pay only compute.
No-code creators who want an appWeakThe open weights need a GPU or ComfyUI; only the capped free GigaChat path is truly no-code.
Long-form or cinematic-resolution videoWeakClips are ~5 seconds at a modest base resolution; Full HD needs the separate upscaler and length is limited.
A finished, captioned, scheduled post from one toolWeakThe model outputs raw footage only — captions, reframing, and scheduling are not included.
Consistent branded persona/avatar channelWeakThere is no persona identity, brand-voice control, or recurring cadence layer; this is a clip generator.

Alternatives worth considering

  • Kompozy — not a rival model but the production and publishing layer: use a Kandinsky clip as B-roll or a hook, then caption, reframe, and schedule across platforms.
  • Alibaba Wan 3.0 — another open-weight video model worth comparing for quality and license terms.
  • LTX-2.5 — a fast open-weight video option if speed and local control matter more than built-in audio.
  • Google Veo 3 — a closed, hosted model if you want higher fidelity and are fine paying per render.
  • Kling AI — a closed model with strong motion and longer clips if open weights are not a requirement.

How Kompozy compares

The cleanest way to place Kompozy here is that it is not an alternative to Kandinsky 6.0 Video — it is the layer that starts where Kandinsky stops. Kandinsky is a model: it hands you a raw 5-second clip, with or without its synchronized audio, and the moment it finishes you are holding footage, not a post. Kompozy is a product: it takes that clip as B-roll or an opening hook, builds the short around it with captions sized for silent autoplay, reframes it per platform, and schedules it across the eight social platforms plus blog and email from one queue. One is a generation primitive; the other is the operation that turns primitives into a published channel.

Honestly, that also means Kompozy generates things Kandinsky does not — face-locked persona and avatar video, carousels, quote graphics, text posts, blogs, and newsletters, all governed by a brand-voice brief and kept on a recurring cadence behind a review gate. If your goal is a self-hosted clip generator you fully control, Kandinsky is the better fit and nothing in Kompozy replaces it. If your goal is a steady, on-brand, multi-platform presence, the model is the cheap input and Kompozy is the engine — the two are complements, not competitors.

Frequently asked questions

Is Kandinsky 6.0 Video free?

Yes. The code and model weights are published under an MIT license and are free to download from Hugging Face, and the same generate-with-sound capability is available at no cost inside Sber's GigaChat assistant. Running the open weights still requires a capable GPU, so you pay compute rather than a per-clip fee.

Can I use Kandinsky 6.0 Video commercially?

The MIT license permits commercial use and self-hosting of both the code and the weights. As always, confirm the exact license terms in the model repository before shipping commercial work, since licensing details can be updated.

How long and what resolution are Kandinsky 6.0 clips?

Base generation produces roughly 5-second clips at 24 fps at a modest resolution. A separate Kandinsky 6.0 super-resolution model upscales the output to Full HD (1920×1080). The clip-length cap inside GigaChat is currently about five seconds, which Sber has said it plans to raise.

What is the difference between Kandinsky 6.0 Pro and Lite?

Pro is the larger 29-billion-parameter line aimed at higher quality; Lite is a 3-billion-parameter line that is lighter to run, making it plausible on a single high-end consumer GPU. Each ships in pretrained and distilled (faster, fewer-step) variants.

Is the synchronized audio any good?

The model generates 44 kHz audio in sync with the video — lip-synced speech, ambient effects, and music — which is the release's headline feature and uncommon in open models. Quality varies by prompt, and you can also generate silent video when you plan to add your own audio.

How does Kandinsky 6.0 compare to Veo or Kling?

Kandinsky 6.0 is open-weight, free, and audio-native but limited to short clips at modest base resolution; Veo and Kling are closed, hosted, paid models that generally offer higher fidelity and longer clips. Direct quality comparisons at this stage are early and often self-reported, so test on your own prompts before deciding.

Can Kandinsky 6.0 Video publish to social platforms?

No. It is a generation model that outputs raw clips — it does not caption, reframe for different feeds, or schedule posts. Turning its output into finished, published content requires a separate editing and distribution step, which is what a content engine like Kompozy provides.

Related deep guides

See Kandinsky 6.0 Video vs Kompozy comparison → · Get Started →