AIAny
Icon for item

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Generates 5-second text- or image-conditioned videos with synchronized 44 kHz audio (including lip-sync) and built-in super-resolution to 1920×1080; available in Lite (3B) and Pro (29B) variants with code and checkpoints released under an MIT license.

Introduction

Multimedia generation is shifting from separately produced visuals and audio toward tightly synchronized audio–video output. Kandinsky 6.0 Video shows that diffusion-based foundation models can generate short, semantically aligned clips by explicitly modeling and aligning an audio stream with a pretrained video stream, producing 5‑second clips with 44 kHz audio and lip‑sync.

Key Findings
  • Dual-stream CrossDiT architecture: connects a pretrained video stream and a newly trained audio stream via bidirectional cross-attention to achieve temporal and semantic alignment across modalities, rather than post-hoc synchronization.
  • Training recipe: continuous pretraining (audio-from-scratch on large audio corpora, then joint audio–video training that preserves unimodal fidelity), followed by supervised fine-tuning, RL-based post-training, and distillation to smaller, faster variants.
  • Outputs and variants: generates 5s clips with synchronized 44 kHz audio in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; offers Lite (≈3B) and Pro (≈29B) models; includes a plug-in super-resolution path to reach Full-HD (1920×1080).
  • Empirical results and release: human evaluations show Pro outperforms Kandinsky 5.0 Video Pro and is competitive with leading audio–video models on speech quality; code, model checkpoints, and diffusers integration are released under the MIT license to accelerate reproducibility.
Who it's for and trade-offs

Great fit if you need research-grade, open-source audio–video generation for short clips (T2AV/I2AV), want access to checkpoints and diffusers integration, and can accept substantial compute for training or inference. Look elsewhere if you require longer videos, low-latency real-time generation, or lightweight on-device models: the system focuses on 5-second clips and uses large diffusion models that demand GPUs and careful memory/offload strategies. Also consider downstream safety and copyright moderation when deploying synthesized speech and video.

Where it fits

Kandinsky 6.0 Video positions itself at the intersection of multimodal foundation models and practical open research: compared with single-modality video generators, it explicitly models audio and enforces alignment during generation; compared with closed commercial multimodal products, it prioritizes openness (MIT) and researcher access to checkpoints and diffusers tooling.

Methodological notes

The bidirectional cross-attention between streams is the core mechanism for alignment; continuous pretraining that first builds unimodal audio fidelity before joint multimodal tuning helps maintain quality in each modality. Distillation and RL-based post-training are used to improve usability and speech naturalness in distilled variants.

Information

  • Websitearxiv.org
  • OrganizationsKandinsky Lab
  • AuthorsTeam Kandinsky, Julia Agafonova, Bulat Akhmatov, Mikhail Aksyutin, Grigorii Alekseenko, Anastasia Aliaskina, Olga Androsova, Vladimir Arkhipkin, Anna Averchenkova, Alexander Belykh …
  • Published date2026/10/04

Categories

More Items

Introduces LoHi, a training-free, single-pass method that mixes dense low-resolution video streams with sparse high-resolution frames to improve long-video vision-language model accuracy under strict token budgets while cutting front-end decoding latency.

Predicts compact 'prospective tokens' that summarize upcoming information needs and uses them to select a small set of past frames for conditioning long-horizon video generation, improving long-range consistency, visual quality, and action alignment while remaining plug-and-play across diverse generators.

Turns live video into reusable textual memory and timely responses by training a streaming video LLM to proactively generate time-grounded captions and event summaries. Key components: Proactive Hierarchical Caption Memory (PHCM) for multi-scale records and Proactive State Transition Learning (PSTL) to balance response timing; trained on the OneStreamer-1M streaming dataset.