Sber's open-weight, MIT-licensed video model family that generates 5-second clips with synchronized audio — lip-synced speech, ambience, and music — in a single pass.
Last verified · 2026-10-10 · by Moe Ameen
Kandinsky 6.0 Video is an open-weight text-to-video and image-to-video model family from Kandinsky Lab, the generative-media team at Russian technology company Sber. It was open-sourced on October 6, 2026 with code and weights under an MIT license, which permits commercial use and self-hosting. The defining feature is audio: the model generates synchronized 44 kHz sound alongside the picture — speech with lip-sync, ambient effects, and music — and can also produce silent video when you plan to add your own.
The release is a family, not a single model. There is a 29-billion-parameter Pro line and a lighter 3-billion-parameter Lite line, each shipping in pretrained and distilled (faster, fewer-step) variants. The models generate roughly 5-second clips at 24 fps at a modest base resolution, and a separate Kandinsky 6.0 super-resolution model upscales the output to Full HD (1920×1080). Because the Lite line is comparatively small, it is plausible to run on a single high-end consumer GPU rather than data-center hardware.
Access is unusually open. Weights live on Hugging Face under the `kandinskylab` organization, code is on GitHub, and the launch arrived with Diffusers pipelines, a ComfyUI extension, vLLM-Omni support, a public demo Space, and Colab/Kaggle notebooks. For non-technical users, the same generate-with-sound capability is available free inside Sber's GigaChat assistant, where the clip-length cap is currently about five seconds.
A note on specifics: parameter counts, resolution details, and the Full HD upscaler come from the model's documentation and launch coverage, and Sber's own quality claims (including that Pro outperforms its predecessor in human side-by-side evaluation) and any comparisons against closed models like Veo or Kling are early and self-reported. Verify current details against the model repository before relying on them.
Kandinsky's sweet spot is cheap, repeatable, sound-equipped short clips you can batch on your own hardware. The thing to understand is what a 5-second clip actually is in a content workflow: it is a hook or a B-roll beat, not a post. That is exactly the slot [Kompozy](/) is built to fill. Generate a handful of Kandinsky clips — a punchy visual with its native audio, or a silent loop you will caption — and bring one in as the opening hook of a Marketing Short or as B-roll inside a [Persona Short](/glossary/persona-shorts), where Kompozy wraps it with a scripted message, burned-in captions sized for silent autoplay, and the correct 9:16 / 1:1 / 16:9 framing for each feed. The raw clip becomes a finished segment instead of an orphaned file sitting in a downloads folder.
The bigger win is everything Kandinsky structurally cannot make. A text-to-video model gives you one short clip; a channel needs variety and a voice. Kompozy generates the formats around it — face-locked persona and avatar video, carousels, quote graphics, persona tweets, text posts, a blog, and a newsletter — all held on-brand by a [Persona Brief](/glossary/persona-brief), then fans and schedules the whole set across the eight social platforms plus blog and email. So the division is clean and complementary: run Kandinsky locally for near-zero-cost clips, and let Kompozy turn each one into reviewed, captioned, scheduled content while generating the rest of the cadence the model never touches.
It is an open-weight AI video model family from Sber's Kandinsky Lab, open-sourced October 6, 2026 under an MIT license. It generates roughly 5-second clips at 24 fps from text or an image and produces synchronized 44 kHz audio — lip-synced speech, ambience, and music — alongside the video, with silent output optional.
Yes. Code and weights are MIT-licensed and free to download from Hugging Face, which also allows commercial use and self-hosting. The same capability is available free inside Sber's GigaChat assistant. Running the open weights requires a capable GPU, so the real cost is compute rather than a per-clip fee.
Base generation produces about 5-second clips at 24 fps at a modest resolution, and a separate Kandinsky 6.0 super-resolution model upscales to Full HD (1920×1080). Inside GigaChat the clip-length cap is currently around five seconds, which Sber has said it plans to increase.
The Pro line is large (29B parameters), but the 3B Lite line is light enough to be plausible on a single high-end consumer GPU. It runs through Hugging Face Diffusers, ComfyUI, and vLLM-Omni, with Colab and Kaggle notebooks available if you do not have local hardware.
The model outputs a short raw clip, so publishing it needs a hook, captions, per-platform reframing, and scheduling. A content engine like Kompozy takes the clip as B-roll or an opening hook, builds the short around it, and schedules it across the eight social platforms plus blog and email.