An honest review of Kandinsky 6.0 Video, Sber's free MIT-licensed AI video model with synchronized audio — scores, pros and cons, and who it's actually for.
As a free, MIT-licensed, open-weight model that generates 5-second clips with synchronized audio, Kandinsky 6.0 Video is a genuinely notable release — audio-native generation you can run yourself and use commercially is rare. But it is a model, not a product: the output is short, modest-resolution, and silent-optional raw footage, and getting published content out of it means your own GPU or ComfyUI, plus captions, reframing, and a schedule. Rate it high as a building block and low as a finished workflow.
On October 6, 2026, Sber's Kandinsky Lab open-sourced Kandinsky 6.0 Video — a family of video models that generate clips with synchronized sound, released with code and weights under an MIT license. There are two lines, a 29-billion-parameter Pro and a 3-billion-parameter Lite, and the same generate-with-sound capability is also live for free inside Sber's GigaChat assistant. For a category where most open models output silent video, shipping lip-synced speech, ambience, and music in the same pass is the headline.
I run a competing content engine, so weigh that, but this review is about the model on its own terms, not a sales pitch. The honest question for a creator is not "is the research impressive?" — it clearly is. It is "what do I actually get, and what do I still have to do myself?" Kandinsky 6.0 Video gives you a free, commercially usable, audio-equipped clip generator. What it does not give you is anything resembling a finished post: the clips are about five seconds, the base resolution is modest (a separate super-resolution model upscales to Full HD), and running the open weights assumes a capable GPU or a ComfyUI setup.
So this page scores two different things that searchers conflate: the model as a generation primitive, where it is strong, and the model as a way to produce publishable content, where it is only a first step. Everything below reflects the release as documented on 2026-10-10; where quality claims against closed models are early or self-reported, I say so.
Kandinsky 6.0 Video is an open-weight text-to-video and image-to-video model family from Kandinsky Lab, the generative-media team at Sber. It generates roughly 5-second clips at 24 fps and produces synchronized 44 kHz audio alongside the picture — covering speech with lip-sync, ambient effects, and music — with the option to generate silent video instead. It ships as a 29B Pro line and a 3B Lite line, each in pretrained and distilled (faster, fewer-step) variants, and a separate Kandinsky 6.0 super-resolution model upscales the modest base resolution to Full HD (1920×1080). Access is unusually open for a model with these capabilities. Weights are on Hugging Face under the `kandinskylab` organization, code is on GitHub, and the release arrived with Diffusers pipelines, a ComfyUI extension, vLLM-Omni support, a demo Space, and Colab/Kaggle notebooks. Because the license is MIT, the weights can be used commercially and self-hosted — you pay compute, not a per-clip fee. For non-technical users, the same capability is available free inside GigaChat, where the clip-length cap is currently about five seconds.
Kandinsky 6.0 Video fits technically comfortable creators and small teams who want cheap, repeatable, sound-equipped short clips and are happy to run a model locally or in ComfyUI — think B-roll, hooks, loops, and experimentation at volume without per-render metering. It is a weak fit for anyone who wants a finished, captioned, correctly-framed, scheduled post out of a single tool, or who needs clips longer than a few seconds, cinematic resolution out of the box, or a no-code path beyond GigaChat's free tier. If you are not going to touch a GPU or a node graph, the open weights are not really for you; the GigaChat path is.
| Dimension | Score | Why |
|---|---|---|
| Generation quality | 4.0 / 5 | Strong for an open 5-second model; Sber's own report says human reviewers preferred Pro over its predecessor, though head-to-head claims against closed models are early. |
| Synchronized audio | 4.5 / 5 | The standout: lip-synced speech, ambience, and music generated with the video in one pass, at 44 kHz — rare in open models. |
| Clip length & resolution | 3.0 / 5 | About 5 seconds at a modest base resolution; Full HD needs the separate upscaler, and longer clips are not yet supported. |
| Openness & licensing | 5.0 / 5 | MIT license on both code and weights, commercial use allowed, self-hostable — as open as it gets for a model this capable. |
| Ecosystem & integrations | 4.5 / 5 | Launched with Diffusers, ComfyUI, vLLM-Omni, a demo Space, and Colab/Kaggle notebooks — unusually complete at release. |
| Ease of use | 2.5 / 5 | Running the open weights assumes a capable GPU or ComfyUI. The free GigaChat path is easy but limited; there is no polished standalone app. |
| Value | 4.5 / 5 | Free weights and a free GigaChat tier make the cost compute-only; for high-volume short clips that is excellent value. |
| Publishing readiness | 2.0 / 5 | Output is raw footage — no captions, no reframing, no scheduling. Getting to a published post is entirely on you. |
There is no price in the usual sense, which is the whole point. Kandinsky 6.0 Video's code and weights are MIT-licensed and free to download, so the only cost of running them is compute — a GPU you own or rent. For a creator generating a high volume of short clips, that is a materially different economic model than a hosted generator that meters per render: once you have the hardware, the marginal cost of another 5-second clip approaches zero. The 3B Lite line lowers the hardware bar enough that a single high-end consumer GPU is plausible, which is what makes the "free" claim real rather than theoretical.
The catch is that free weights are not free workflow. The time and skill to stand up Diffusers or ComfyUI, run the separate upscaler for Full HD, and then caption, reframe, and schedule the output are real costs that a per-clip hosted price would otherwise absorb. For non-technical creators, the free GigaChat tier removes the setup cost but adds limits — roughly five-second clips inside Sber's assistant rather than a dedicated creator surface.
Net: as a generation primitive, the value is excellent, especially at volume. As a path to finished, published content, "free" understates the real cost, because the model deliberately stops at the clip. Budget for the production and distribution layer separately — whether that is your own editing time or a tool that does it.
| Use case | Fit | Why |
|---|---|---|
| High-volume short B-roll and hooks | Strong | Free, self-hostable, audio-equipped 5-second clips are ideal raw material to generate at volume. |
| Adding synchronized sound to AI video | Strong | Lip-synced speech, ambience, and music in one pass is exactly what this model is built for. |
| Commercial use on a budget | Strong | The MIT license permits commercial use with no per-clip fee; you pay only compute. |
| No-code creators who want an app | Weak | The open weights need a GPU or ComfyUI; only the capped free GigaChat path is truly no-code. |
| Long-form or cinematic-resolution video | Weak | Clips are ~5 seconds at a modest base resolution; Full HD needs the separate upscaler and length is limited. |
| A finished, captioned, scheduled post from one tool | Weak | The model outputs raw footage only — captions, reframing, and scheduling are not included. |
| Consistent branded persona/avatar channel | Weak | There is no persona identity, brand-voice control, or recurring cadence layer; this is a clip generator. |
The cleanest way to place Kompozy here is that it is not an alternative to Kandinsky 6.0 Video — it is the layer that starts where Kandinsky stops. Kandinsky is a model: it hands you a raw 5-second clip, with or without its synchronized audio, and the moment it finishes you are holding footage, not a post. Kompozy is a product: it takes that clip as B-roll or an opening hook, builds the short around it with captions sized for silent autoplay, reframes it per platform, and schedules it across the eight social platforms plus blog and email from one queue. One is a generation primitive; the other is the operation that turns primitives into a published channel.
Honestly, that also means Kompozy generates things Kandinsky does not — face-locked persona and avatar video, carousels, quote graphics, text posts, blogs, and newsletters, all governed by a brand-voice brief and kept on a recurring cadence behind a review gate. If your goal is a self-hosted clip generator you fully control, Kandinsky is the better fit and nothing in Kompozy replaces it. If your goal is a steady, on-brand, multi-platform presence, the model is the cheap input and Kompozy is the engine — the two are complements, not competitors.
Yes. The code and model weights are published under an MIT license and are free to download from Hugging Face, and the same generate-with-sound capability is available at no cost inside Sber's GigaChat assistant. Running the open weights still requires a capable GPU, so you pay compute rather than a per-clip fee.
The MIT license permits commercial use and self-hosting of both the code and the weights. As always, confirm the exact license terms in the model repository before shipping commercial work, since licensing details can be updated.
Base generation produces roughly 5-second clips at 24 fps at a modest resolution. A separate Kandinsky 6.0 super-resolution model upscales the output to Full HD (1920×1080). The clip-length cap inside GigaChat is currently about five seconds, which Sber has said it plans to raise.
Pro is the larger 29-billion-parameter line aimed at higher quality; Lite is a 3-billion-parameter line that is lighter to run, making it plausible on a single high-end consumer GPU. Each ships in pretrained and distilled (faster, fewer-step) variants.
The model generates 44 kHz audio in sync with the video — lip-synced speech, ambient effects, and music — which is the release's headline feature and uncommon in open models. Quality varies by prompt, and you can also generate silent video when you plan to add your own audio.
Kandinsky 6.0 is open-weight, free, and audio-native but limited to short clips at modest base resolution; Veo and Kling are closed, hosted, paid models that generally offer higher fidelity and longer clips. Direct quality comparisons at this stage are early and often self-reported, so test on your own prompts before deciding.
No. It is a generation model that outputs raw clips — it does not caption, reframe for different feeds, or schedule posts. Turning its output into finished, published content requires a separate editing and distribution step, which is what a content engine like Kompozy provides.
See Kandinsky 6.0 Video vs Kompozy comparison → · Get Started →