AIAny
Icon for item

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Trains a single unified multimodal model with reinforcement learning to perform end-to-end self-reflection and iterative image repair — jointly learning the diagnostic (textual) reflection and the flow-based image revisions so credit propagates across rounds without an external verifier.

Introduction

Most image editing or multi-round generation pipelines split diagnosis, planning, and rendering across separate modules or rely on supervised traces for a cold start. This paper argues that unified multimodal models can learn native, interleaved reflection — generating diagnostic text, rendering edits, observing the result, and continuing the loop — but only if training assigns credit across the whole trajectory rather than per-round or per-module.

Key Findings
  • Joint trajectory-level RL: Introduces group-relative and trajectory-level advantage computations so sibling rollouts (sharing the same initial image) can be compared and a single advantage updates both reflection tokens and flow-based image revisions, avoiding combinatorial credit assignment.
  • Removes external verifier at inference: Because credit flows across rounds and to both roles of the same model, no separate critic or verifier is required at test time; the model learns to detect and fix its own mistakes natively.
  • Empirical gains and generalization: Training with UMM-Reflection yields substantial improvements on held-out benchmarks, e.g., +12.05 GenEval on BAGEL and transferable gains on WISE (+10.97), OneIG-Bench (+3.48) and T2I-CompBench++ (+4.63), despite none of these being used in training.
  • Practical design choices: Uses sibling-group comparisons to stabilize RL signals and updates both text- and image-headed parameters together, which the authors show outperforms optimizing only the renderer or only the reflection head.
Who it's for and trade-offs

Great fit if you are researching multimodal models that should autonomously diagnose and iteratively repair visual outputs, or if you need a single-model solution that avoids external critics at inference. Expect improved cross-benchmark transfer for iterative editing and reasoning tasks. Look elsewhere if you need a lightweight plug-in editor or a strict modular pipeline with separately maintainable verifier components: joint trajectory RL increases training complexity, requires rollout grouping and careful reward design, and may demand more compute and engineering to stabilize than conventional SFT or single-module RL baselines.

Information

  • Websitearxiv.org
  • AuthorsYijia Fan, Ziqi Huang, Zhongang Cai, Yan Li, Zimo Wen, Wanqi Yin, Haiwen Diao, Ziwei Liu
  • Published date2026/09/28

More Items

Normalizes each domain's teacher feedback spread during multi-teacher on-policy distillation so that no domain (e.g., instruction following) overwhelms others, improving student recovery of specialist skills and raising average scores across benchmarks.

Uses diffusion-model forking moments as a proxy for perceptual distance to automatically generate pointwise reference-grounded labels, enabling annotation-free training of reference-based image quality assessment metrics.

Learns joint predictive visual dynamics and action generation for generalist robot manipulation, translating future-relevant visual representations into actions. Integrates a Mixture-of-Transformers coupling a video expert and action expert, a frozen vision-language model for semantics, 4D distillation, and Causal Imprint; pretrained on a 20K+ hour heterogeneous corpus.