// AI TOOLS · KANDINSKY 6.0 VIDEO

Kandinsky 6.0 Video

Sber's open-weight, MIT-licensed video model family that generates 5-second clips with synchronized audio — lip-synced speech, ambience, and music — in a single pass.

Last verified · 2026-10-10 · by Moe Ameen

What Kandinsky 6.0 Video is

Kandinsky 6.0 Video is an open-weight text-to-video and image-to-video model family from Kandinsky Lab, the generative-media team at Russian technology company Sber. It was open-sourced on October 6, 2026 with code and weights under an MIT license, which permits commercial use and self-hosting. The defining feature is audio: the model generates synchronized 44 kHz sound alongside the picture — speech with lip-sync, ambient effects, and music — and can also produce silent video when you plan to add your own.

The release is a family, not a single model. There is a 29-billion-parameter Pro line and a lighter 3-billion-parameter Lite line, each shipping in pretrained and distilled (faster, fewer-step) variants. The models generate roughly 5-second clips at 24 fps at a modest base resolution, and a separate Kandinsky 6.0 super-resolution model upscales the output to Full HD (1920×1080). Because the Lite line is comparatively small, it is plausible to run on a single high-end consumer GPU rather than data-center hardware.

Access is unusually open. Weights live on Hugging Face under the `kandinskylab` organization, code is on GitHub, and the launch arrived with Diffusers pipelines, a ComfyUI extension, vLLM-Omni support, a public demo Space, and Colab/Kaggle notebooks. For non-technical users, the same generate-with-sound capability is available free inside Sber's GigaChat assistant, where the clip-length cap is currently about five seconds.

A note on specifics: parameter counts, resolution details, and the Full HD upscaler come from the model's documentation and launch coverage, and Sber's own quality claims (including that Pro outperforms its predecessor in human side-by-side evaluation) and any comparisons against closed models like Veo or Kling are early and self-reported. Verify current details against the model repository before relying on them.

What you can make with it

  • Short (~5-second) text-to-video clips with synchronized audio — speech, ambience, or music
  • Image-to-video clips that animate a still with matching sound
  • Silent short clips when you want to add your own audio later
  • Scroll-stopping opening hooks and ambient B-roll generated at volume on your own GPU
  • Full HD (1920×1080) versions of any clip via the separate super-resolution model
  • Lip-synced talking moments for a character, within the ~5-second length limit

How Kompozy turns Kandinsky 6.0 Video output into content

Kandinsky's sweet spot is cheap, repeatable, sound-equipped short clips you can batch on your own hardware. The thing to understand is what a 5-second clip actually is in a content workflow: it is a hook or a B-roll beat, not a post. That is exactly the slot [Kompozy](/) is built to fill. Generate a handful of Kandinsky clips — a punchy visual with its native audio, or a silent loop you will caption — and bring one in as the opening hook of a Marketing Short or as B-roll inside a [Persona Short](/glossary/persona-shorts), where Kompozy wraps it with a scripted message, burned-in captions sized for silent autoplay, and the correct 9:16 / 1:1 / 16:9 framing for each feed. The raw clip becomes a finished segment instead of an orphaned file sitting in a downloads folder.

The bigger win is everything Kandinsky structurally cannot make. A text-to-video model gives you one short clip; a channel needs variety and a voice. Kompozy generates the formats around it — face-locked persona and avatar video, carousels, quote graphics, persona tweets, text posts, a blog, and a newsletter — all held on-brand by a [Persona Brief](/glossary/persona-brief), then fans and schedules the whole set across the eight social platforms plus blog and email. So the division is clean and complementary: run Kandinsky locally for near-zero-cost clips, and let Kompozy turn each one into reviewed, captioned, scheduled content while generating the rest of the cadence the model never touches.

  1. Generate your clip in Kandinsky 6.0 — self-hosted via Diffusers/ComfyUI on your GPU, or free in GigaChat — and upscale to Full HD with the super-resolution model if you need it.
  2. Decide its role: an opening hook for a Marketing Short, or ambient B-roll for a Persona Short. Keep or drop the model's synchronized audio depending on the fit.
  3. In Kompozy, bring the clip into the matching video format and let the engine add the scripted message, burned-in captions, and per-platform aspect ratios.
  4. Generate the surrounding cadence — persona video, carousels, quote cards, a text post, a blog, a newsletter — on-brand via your Persona Brief, so one clip anchors a full week of content.
  5. Review each post in the pipeline, then schedule the set across the eight social platforms plus blog and email from one queue.

Frequently asked questions

What is Kandinsky 6.0 Video?

It is an open-weight AI video model family from Sber's Kandinsky Lab, open-sourced October 6, 2026 under an MIT license. It generates roughly 5-second clips at 24 fps from text or an image and produces synchronized 44 kHz audio — lip-synced speech, ambience, and music — alongside the video, with silent output optional.

Is Kandinsky 6.0 Video free to use?

Yes. Code and weights are MIT-licensed and free to download from Hugging Face, which also allows commercial use and self-hosting. The same capability is available free inside Sber's GigaChat assistant. Running the open weights requires a capable GPU, so the real cost is compute rather than a per-clip fee.

How long are Kandinsky 6.0 clips and what resolution?

Base generation produces about 5-second clips at 24 fps at a modest resolution, and a separate Kandinsky 6.0 super-resolution model upscales to Full HD (1920×1080). Inside GigaChat the clip-length cap is currently around five seconds, which Sber has said it plans to increase.

Can I run Kandinsky 6.0 Video on my own computer?

The Pro line is large (29B parameters), but the 3B Lite line is light enough to be plausible on a single high-end consumer GPU. It runs through Hugging Face Diffusers, ComfyUI, and vLLM-Omni, with Colab and Kaggle notebooks available if you do not have local hardware.

How do I turn a Kandinsky clip into a social post?

The model outputs a short raw clip, so publishing it needs a hook, captions, per-platform reframing, and scheduling. A content engine like Kompozy takes the clip as B-roll or an opening hook, builds the short around it, and schedules it across the eight social platforms plus blog and email.

Related tools

  • Wan3.0 — Alibaba's latest AI video model in the Tongyi Wanxiang (Wan) line. It generates clips up to about 30 seconds — roughly double the prior Wan 2.x generation — from text, images, or documents like PDFs and slide decks. Fully launched August 24, 2026.
  • LTX-2.5 — LTX's open-weight video model — spun out of Lightricks — that turns an image into a 10-second clip in seconds, and runs on a GPU you already own.
  • Google Veo 3 — Google DeepMind's video model that was the first to generate synchronized native audio — dialogue, sound effects, and music — inside the same pass as the video, with lip sync.
  • Kling AI — Kuaishou's text-to-video and image-to-video model — turn a prompt or a still into a cinematic clip with camera motion, lip sync, and native audio.
  • MiniMax H3 Video Model — MiniMax's open-weights, multimodal video model — 2K clips with native stereo audio, conditioned on up to 9 image, 3 video, and 3 audio references, priced to undercut the proprietary leaders.
  • ByteDance Seedance 2.5 — AI video model that generates a 30-second clip in one pass — no stitching.

← All AI tools · Get started →