AIAny
Icon for item

DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation

Combines joint distribution distillation from a video teacher with marginal (frame-level) distillation from an image teacher to improve few-step video generation. Introduces LatentBridge to align incompatible latents and Latent Variation Sampling to distribute frame supervision, boosting per-frame visual quality and semantic alignment while largely preserving motion.

Introduction

Most few-step video generators face a fundamental trade-off: joint sequence-level supervision preserves temporal dynamics but can leave frame-level appearance and semantic alignment weak; frame-level image priors improve single-frame realism but can disrupt motion. DuoMatching's core insight is pragmatic: approximate the true video distribution with a unified joint–marginal objective that leverages both a video teacher (joint DMD) and an image teacher (marginal DMD) so each source compensates the other's weaknesses.

Key Findings
  • Joint–marginal objective: A weighted combination of joint DMD (video teacher) and marginal DMD (image teacher) steers the student toward better temporal consistency and stronger frame-level appearance/semantic priors, respectively, with a tunable weight ω controlling the balance.
  • LatentBridge: A lightweight adapter that reconciles incompatible latent spaces between temporally compressed video latents and image-model latents, enabling effective frame-wise supervision without overly suppressing encoded dynamics.
  • Latent Variation Sampling (LVS): Samples frame marginals across distinct temporal segments to reduce redundant supervision and better preserve motion diversity while still transferring image priors.
  • Empirical gains: Improves fine-grained visual quality, compositional coherence, and prompt-specified semantic alignment in few-step video generation; human evaluations report >80% overall preference versus evaluated baselines while largely maintaining temporal quality.
Who it's for and trade-offs

Great fit if you train or distill few-step/streaming video generators and need better per-frame realism or stronger prompt alignment without fully sacrificing motion dynamics. It is especially useful when you can access both a video-level teacher and a high-quality image generator to provide complementary priors.

Look elsewhere if your priority is extreme long-horizon temporal fidelity beyond the model's training rollout regime (DuoMatching targets few-step streaming scenarios) or if you cannot align latent representations between student and image teacher (LatentBridge is proposed to help, but some engineering is required).

Method highlights
  • Training objective: J_DM(θ)=D_KL(Q_θ ∥ P_v) + ω D_KL(m_θ ∥ P_i), where Q_θ is the student joint, P_v the video-teacher joint, m_θ the frame marginal from the student, and P_i the image-teacher marginal. ω trades joint vs. marginal influence.
  • Implementation notes: Apply marginal DMD to a sampled subset of frames per video; use LatentBridge to map video latents into a representation compatible with the image teacher; apply LVS to pick frames from different temporal segments to avoid redundant supervision.

Overall, DuoMatching reframes distillation for generative video by explicitly mixing joint sequence constraints with dedicated frame-level guidance, yielding clearer frames and better semantic alignment while retaining temporal behaviors.

Information

  • Websitearxiv.org
  • AuthorsJiahao Zhan, Yan Wang, Yongrui Ma, Qunliang Xing, Ruchang Yao, Runtao Liu, Shijie Zhao, Tianfan Xue
  • Published date2026/10/02

More Items

Introduces LoHi, a training-free, single-pass method that mixes dense low-resolution video streams with sparse high-resolution frames to improve long-video vision-language model accuracy under strict token budgets while cutting front-end decoding latency.

Predicts compact 'prospective tokens' that summarize upcoming information needs and uses them to select a small set of past frames for conditioning long-horizon video generation, improving long-range consistency, visual quality, and action alignment while remaining plug-and-play across diverse generators.

Lets general-purpose vision-language models directly command robots via a compact mid-level action interface and asynchronous monitoring, enabling zero-shot manipulation without task-specific policy training; demonstrates strong sim benchmarks and real xArm6 transfer.