Multimedia generation is shifting from separately produced visuals and audio toward tightly synchronized audio–video output. Kandinsky 6.0 Video shows that diffusion-based foundation models can generate short, semantically aligned clips by explicitly modeling and aligning an audio stream with a pretrained video stream, producing 5‑second clips with 44 kHz audio and lip‑sync.
Key Findings
- Dual-stream CrossDiT architecture: connects a pretrained video stream and a newly trained audio stream via bidirectional cross-attention to achieve temporal and semantic alignment across modalities, rather than post-hoc synchronization.
- Training recipe: continuous pretraining (audio-from-scratch on large audio corpora, then joint audio–video training that preserves unimodal fidelity), followed by supervised fine-tuning, RL-based post-training, and distillation to smaller, faster variants.
- Outputs and variants: generates 5s clips with synchronized 44 kHz audio in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; offers Lite (≈3B) and Pro (≈29B) models; includes a plug-in super-resolution path to reach Full-HD (1920×1080).
- Empirical results and release: human evaluations show Pro outperforms Kandinsky 5.0 Video Pro and is competitive with leading audio–video models on speech quality; code, model checkpoints, and diffusers integration are released under the MIT license to accelerate reproducibility.
Who it's for and trade-offs
Great fit if you need research-grade, open-source audio–video generation for short clips (T2AV/I2AV), want access to checkpoints and diffusers integration, and can accept substantial compute for training or inference. Look elsewhere if you require longer videos, low-latency real-time generation, or lightweight on-device models: the system focuses on 5-second clips and uses large diffusion models that demand GPUs and careful memory/offload strategies. Also consider downstream safety and copyright moderation when deploying synthesized speech and video.
Where it fits
Kandinsky 6.0 Video positions itself at the intersection of multimodal foundation models and practical open research: compared with single-modality video generators, it explicitly models audio and enforces alignment during generation; compared with closed commercial multimodal products, it prioritizes openness (MIT) and researcher access to checkpoints and diffusers tooling.
Methodological notes
The bidirectional cross-attention between streams is the core mechanism for alignment; continuous pretraining that first builds unimodal audio fidelity before joint multimodal tuning helps maintain quality in each modality. Distillation and RL-based post-training are used to improve usability and speech naturalness in distilled variants.