Most post-training methods for joint audio–video generation assume fixed ways to route reward-driven updates and fixed reward weights, but cross-modal functions and reward conflicts evolve during fine-tuning. The core insight of this paper is that adapting both WHERE updates act (token/layer localization) and HOW rewards are combined (modality-aware reweighting anchored by user preferences) yields more stable and effective multi-objective RL for audio–video diffusion models.
Key Findings
- Cross-Modal Influence-Guided Routing: bidirectional cross-attention responses serve as a lightweight proxy to identify which tokens and cross-attention layers are functionally important; token-aware loss reweighting and layer scaling concentrate updates along influential cross-modal pathways so gradient flow is preserved where it matters.
- Preference-Preserving Modality-Aware Reweighting: after a warm-up, the method computes branch-specific reward-gradient interactions and applies them as residual corrections to predefined preference weights, resolving emergent conflicts without discarding user priors or letting strong rewards suppress weaker but necessary objectives.
- Empirical gains: when applied to LTX-2 backbones, the combined approach improves video/audio quality, text–audio alignment, and audio–video synchronization versus strong RL baselines; ablations show token-level routing, layer routing, and residual reweighting each contribute.
Who it's for and trade-offs
Great fit if you work on multimodal generative models that couple streams via cross-attention and need to fine-tune for several quality and alignment metrics at once (e.g., audio/video quality, cross-modal semantic alignment, temporal sync). Look elsewhere if you need a plug-and-play single-reward tuning recipe or if your model has no explicit cross-modal attention structure—this approach relies on access to cross-attention responses and gradient interactions and adds complexity compared with fixed-weight scalar aggregation.
Where it fits
This sits between scalar weighted-sum reward tuning and gradient-space multi-reward methods: it localizes updates along cross-modal interaction paths like routing-based methods, while keeping user-defined preference priors and only applying learned residual corrections to reward weights, aiming to balance principled adaptation with predictable priorities.
Mechanism notes
The routing uses layer- and token-level aggregations of cross-attention response norms to derive token weights and layer scales; reweighting estimates conflicts within modality branches and applies residual corrections after warm-up to adapt over training. Reported probes show the attention-response proxy correlates strongly with true functional influence and that static routing becomes stale as fine-tuning progresses.