Kandinsky 6.0 Video is a new lineup of models that generate video with synchronized sound, with the code and model weights released under the MIT license: anyone can use the models free of charge and build their own products on top of them. The lineup includes two models — the flagship Kandinsky 6.0 Video Pro with 29 billion parameters and the lightweight Kandinsky 6.0 Video Lite with 3 billion parameters. Both create videos with dialogue, music, and ambient sound: the audio track syncs with what happens in the frame, and the characters' lip movements match their speech. To get a finished clip with voiceover, you can describe a scene in text or upload the first frame and add a description of the audio track. The maximum duration of a clip with sound is five seconds, and the developers plan to increase it over time.
Model capabilities
Model capabilities AI Generated
Kandinsky 6.0 Video is a multimodal model that simultaneously 'sees' and 'hears' what it generates. As a result, dialogue, effects, and background music are synchronized with the picture, and the character's lips move in sync with speech. If sound is not needed, you can ask the neural network to create a video without it.
Kandinsky produces natural sound across a wide range: from speech and digital effects to footsteps, wind noise, or the sound of a passing car. The model produces clean, detailed sound without 'hissing' or quality loss: 44 kHz, the standard quality of most streaming services.
These capabilities come in handy for a wide variety of tasks: you can bring a family photo to life, record a video greeting with speech or music for a loved one, shoot an ASMR video with the rustle of paper, the crackle of firewood, or the crunch of snow, or animate an old meme or a movie frame.
The new models are available to developers as open source under the MIT license — they can be integrated into their own products free of charge. The model weights are published on Hugging Face https://huggingface.co/kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers, and the code — on GitHub https://github.com/kandinskylab/kandinsky-6/. Kandinsky 6.0 Video works out of the box in popular tools for launching and integrating neural networks: Diffusers, ComfyUI, FastVideo, SGLang, and vLLM-omni. Developers don't need to write integrations from scratch: Kandinsky 6.0 Video can be plugged directly into an existing product. A detailed technical report covering the architecture, data, training pipeline, and all side-by-side measurements is available at https://huggingface.co/papers/2610.05608.
The lineup
Kandinsky 6.0 Video Pro is the lineup's flagship. With 29 billion parameters, it delivers maximum generation quality: five-second clips with 44 kHz sound and lip-sync, created from text or an image, with a built-in super-resolution module raising the picture to Full HD. In side-by-side human comparisons, the Pro model clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, across all evaluation criteria. It beats the open-source LTX 2.5 in visual quality and speech quality, and among open models it trails only MiniMax H3. In the image animation task, Kandinsky 6.0 Video Pro surpasses the leading closed model Veo 3.1 Fast — in visual quality, adherence to the source image, and camera movement.
Kandinsky 6.0 Video Lite is a lightweight version with 3 billion parameters. It shares the flagship's architecture and capabilities — the same five-second clips with synchronized sound, lip-sync, and upscaling to Full HD — but is adapted to run on consumer NVIDIA graphics cards, from the RTX 5060 Ti to the RTX 5090, with ready-made presets for 16, 24, and 32 GB of video memory. Thanks to block offloading, which streams the model's weights through the GPU in parts, both models can be launched even on a home computer.
Video quality
Video quality clipfly.ai
Both models generate clips up to Full HD: the base generation runs in SD, and the built-in super-resolution module upscales the image to HD or Full HD, so you can balance speed and detail. Kandinsky 6.0 Video more accurately conveys real-world physics. The movements of people, animals, and objects in the clips look more natural, and entire scenes look more realistic. This is especially noticeable in dynamic clips with fast motion or camera angle changes.
How the model was trained
Kandinsky 6.0 Video consists of two connected modules that generate video and sound. The audio module was first trained from scratch on 40 mn audio tracks and then trained jointly with the video module on 7 million paired audio-video segments. For training, the team selected only clips with clean, high-quality sound, as well as videos with close-up shots of speaking people for more precise synchronization of speech and facial expressions. Reinforcement-learning post-training reduced the word error rate in generated speech by 47% for the Pro model.
The previous version of Kandinsky spent more than six months in first place among open models in the text-to-video category on arena.ai
From Your Site Articles
Related Articles Around the Web