AIAny
Icon for item

Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL

Proposes a forward-process RL method that dynamically localizes gradient updates and adapts multi-reward coordination for joint audio–video diffusion models. Uses bidirectional cross-attention responses for token/layer routing and preference-preserving reweighting to balance competing objectives, improving modality quality, alignment, and synchronization.

Introduction

Most post-training methods for joint audio–video generation assume fixed ways to route reward-driven updates and fixed reward weights, but cross-modal functions and reward conflicts evolve during fine-tuning. The core insight of this paper is that adapting both WHERE updates act (token/layer localization) and HOW rewards are combined (modality-aware reweighting anchored by user preferences) yields more stable and effective multi-objective RL for audio–video diffusion models.

Key Findings
  • Cross-Modal Influence-Guided Routing: bidirectional cross-attention responses serve as a lightweight proxy to identify which tokens and cross-attention layers are functionally important; token-aware loss reweighting and layer scaling concentrate updates along influential cross-modal pathways so gradient flow is preserved where it matters.
  • Preference-Preserving Modality-Aware Reweighting: after a warm-up, the method computes branch-specific reward-gradient interactions and applies them as residual corrections to predefined preference weights, resolving emergent conflicts without discarding user priors or letting strong rewards suppress weaker but necessary objectives.
  • Empirical gains: when applied to LTX-2 backbones, the combined approach improves video/audio quality, text–audio alignment, and audio–video synchronization versus strong RL baselines; ablations show token-level routing, layer routing, and residual reweighting each contribute.
Who it's for and trade-offs

Great fit if you work on multimodal generative models that couple streams via cross-attention and need to fine-tune for several quality and alignment metrics at once (e.g., audio/video quality, cross-modal semantic alignment, temporal sync). Look elsewhere if you need a plug-and-play single-reward tuning recipe or if your model has no explicit cross-modal attention structure—this approach relies on access to cross-attention responses and gradient interactions and adds complexity compared with fixed-weight scalar aggregation.

Where it fits

This sits between scalar weighted-sum reward tuning and gradient-space multi-reward methods: it localizes updates along cross-modal interaction paths like routing-based methods, while keeping user-defined preference priors and only applying learned residual corrections to reward weights, aiming to balance principled adaptation with predictable priorities.

Mechanism notes

The routing uses layer- and token-level aggregations of cross-attention response norms to derive token weights and layer scales; reweighting estimates conflicts within modality branches and applies residual corrections after warm-up to adapt over training. Reported probes show the attention-response proxy correlates strongly with true functional influence and that static routing becomes stale as fine-tuning progresses.

Information

  • Websitearxiv.org
  • OrganizationsMMLab@HKUST, The Hong Kong University of Science and Technology, Tencent Video, The University of Hong Kong
  • AuthorsSonglin Yang, Xiaotong Zhao, Jiacheng Zhang, Zhe Wang, Toyota Li, Eric Liu, Alan Zhao, Anyi Rao
  • Published date2026/09/29

Categories

More Items

Groups visually grounded appearances of the same physical instance into persistent, retrievable “biographies” so agents can follow objects across hours or days for long-video question answering. Links identity-aware observations to episodic context and visual evidence; improves EgoLifeQA to 72.0% and increases evidence-window reach from 37.6% to 58.9%.

Transforms user prompts into shot-level cinematic directions for text-to-video generation, using a 397B prompt-enhancer trained on 1.05M videos and SC-GRPO to preserve semantic consistency across shots; evaluated on WanPEval (5–30s) with large human-preference gains.

Provides WROP: a 1.5M-sample synthetic video corpus and a 300-question exam for training and evaluating object permanence and solidity in video world models. Includes 150 Blender task generators, a human Elo benchmark across 14 models, and a fine-tuned 16B continuation model (PWM-WROP).