Kandinsky Lab · 2026-10-04 · notable
Kandinsky 6.0 Video — open MIT models make video with synced speech and sound
Kandinsky 6.0 Video is an MIT-licensed family of 3B and 29B diffusion models that make 5-second clips with synced 44 kHz audio and lip-sync. A 1.4B super-resolution model lifts output to Full HD.

An open video model family that generates picture and sound together, with lip-sync, under the MIT license.
What is it?
Kandinsky 6.0 Video adds synchronized sound to Kandinsky Lab's video models. Kandinsky 6.0 Video Lite (3B) and Pro (29B) make 5-second clips with 44 kHz audio, including lip-synced speech, from a text prompt or from an image plus text. A separate 1.4B Kandinsky 6.0 VSR model raises the output to Full HD (1920×1080).
How does it work?
The models use a dual-stream CrossDiT: a pretrained video stream and a newly trained audio stream are joined by bidirectional cross-attention so picture and sound line up in time and meaning. The audio stream is first trained alone on large audio corpora, then both streams train together on paired audio-video data, followed by supervised fine-tuning, reinforcement-learning post-training and distillation.
Why does it matter?
Open audio-video generation is rare, and this release ships code, checkpoints and a diffusers integration under the permissive MIT license. In the team's side-by-side human evaluation, the Pro model clearly beats Kandinsky 5.0 Video Pro and stays competitive with leading audio-video models, especially on speech quality, the paper says.
Who is it for?
video-generation researchers, open-model and ComfyUI users
Try it
git clone https://github.com/kandinskylab/kandinsky-6.git && just setup && just download pro-distill