Sber's Kandinsky Lab released two open-weight video models — a 29B Pro and a 3B Lite — under an MIT license. Both generate 5-second clips at 24 fps with synchronized 44 kHz audio, including lip-sync, and a separate super-resolution model upscales the output to Full HD. The models are on Hugging Face and GitHub, and the capability is also live free inside Sber's GigaChat assistant.
2026-10-10 · by Moe Ameen
Kandinsky Lab, the generative-media team at Russian technology company Sber, open-sourced Kandinsky 6.0 Video on October 6, 2026. The release ships both code and model weights under an MIT license — the permissive license that allows commercial use — and the weights are hosted on Hugging Face under the `kandinskylab` organization, with code on GitHub at `kandinskylab/kandinsky-6` and an accompanying technical report on arXiv. The same capability is also available for free inside Sber's GigaChat assistant, so a non-technical user can generate clips without touching the open weights.
The release is a family rather than a single model. There is a 29-billion-parameter "Pro" line and a lighter 3-billion-parameter "Lite" line, and each ships in pretrained and distilled (fewer-step, faster) variants. The models generate roughly 5-second clips at 24 fps and support both text-to-video and image-to-video. The headline feature is sound: the models produce synchronized 44 kHz audio alongside the picture — covering speech with lip-sync, ambient effects, and music — and can also generate silent video when audio is not wanted. Base generation runs at a modest resolution; a separate Kandinsky 6.0 super-resolution model upscales the result to Full HD (1920×1080).
Beyond raw Diffusers usage, the launch arrived with integrations already in place — the models run through Hugging Face Diffusers pipelines, vLLM-Omni, and a ComfyUI extension, with a public demo Space and notebooks for Google Colab and Kaggle. The accompanying technical report says human reviewers preferred Pro's output over its predecessor, Kandinsky 5.0 Video Pro, in side-by-side comparisons. A note on the specifics: parameter counts, the resolution details, and the Full HD upscaler come from the model's own documentation and launch coverage; Sber's clip-length cap inside GigaChat is currently about five seconds, which the company has said it plans to raise. Benchmark comparisons against closed models such as Veo or Kling are early and self-reported, so treat head-to-head quality claims as unconfirmed.
If you want to act on this today, the fastest move is to treat a Kandinsky 6.0 clip as a raw ingredient and let Kompozy do the finishing and distribution. A 5-second generated clip — with or without the model's synchronized audio — is exactly the kind of asset [Kompozy](/) is built to turn into published content: drop it in as the opening hook or B-roll, and generate the surrounding short around it with burned-in captions sized for silent autoplay, then schedule the result across the eight social platforms plus blog and email from one queue. You are not hand-editing a bare clip in a timeline; you are feeding it into an engine that adds the hook, the captions, the per-platform reframing, and the posting.
The deeper point is that one 5-second clip is a single post, and a channel needs a cadence. Kompozy generates the formats a text-to-video model does not — face-locked persona shorts and longer avatar video, carousels, quote graphics, text posts, a blog, and a newsletter — all governed by a [Persona Brief](/glossary/persona-brief) that keeps the voice consistent, and [Autopilot](/glossary/autopilot) keeps that mix running on a recurring schedule behind a per-post review gate. So Kandinsky gives you cheap, open, sound-equipped clips; Kompozy wraps them into a steady, on-brand, multi-platform presence. For other open-weight video releases worth watching, see [Alibaba's Wan 3.0](/news/alibaba-wan-3-0-launch), [LTX-2.5](/news/ltx-2-5-launch), and [MiniMax's open-weights video model](/news/minimax-h3-open-weights-video-model-launch).
It is an open-source family of AI video models from Sber's Kandinsky Lab, released October 6, 2026 under an MIT license. It generates roughly 5-second clips at 24 fps from text or an image, and its defining feature is synchronized 44 kHz audio — including lip-synced speech, ambience, and music — produced alongside the picture.
Yes. The code and weights are published under an MIT license, which permits commercial use, and the models are free to download from Hugging Face. The same capability is also available at no cost inside Sber's GigaChat assistant. Running the open weights still costs compute, since you need a suitable GPU.
Base generation produces roughly 5-second clips at 24 fps at a modest resolution, and a separate Kandinsky 6.0 super-resolution model upscales the result to Full HD (1920×1080). The clip-length cap inside GigaChat is currently about five seconds, which Sber has said it plans to increase.
The model gives you a short raw clip; publishing it still needs a hook, captions sized for silent autoplay, per-platform reframing, and scheduling. A content engine like Kompozy takes the clip as B-roll or an opening hook, builds the short around it, and schedules it across the eight social platforms plus blog and email from one queue.