Most image editing or multi-round generation pipelines split diagnosis, planning, and rendering across separate modules or rely on supervised traces for a cold start. This paper argues that unified multimodal models can learn native, interleaved reflection — generating diagnostic text, rendering edits, observing the result, and continuing the loop — but only if training assigns credit across the whole trajectory rather than per-round or per-module.
Key Findings
- Joint trajectory-level RL: Introduces group-relative and trajectory-level advantage computations so sibling rollouts (sharing the same initial image) can be compared and a single advantage updates both reflection tokens and flow-based image revisions, avoiding combinatorial credit assignment.
- Removes external verifier at inference: Because credit flows across rounds and to both roles of the same model, no separate critic or verifier is required at test time; the model learns to detect and fix its own mistakes natively.
- Empirical gains and generalization: Training with UMM-Reflection yields substantial improvements on held-out benchmarks, e.g., +12.05 GenEval on BAGEL and transferable gains on WISE (+10.97), OneIG-Bench (+3.48) and T2I-CompBench++ (+4.63), despite none of these being used in training.
- Practical design choices: Uses sibling-group comparisons to stabilize RL signals and updates both text- and image-headed parameters together, which the authors show outperforms optimizing only the renderer or only the reflection head.
Who it's for and trade-offs
Great fit if you are researching multimodal models that should autonomously diagnose and iteratively repair visual outputs, or if you need a single-model solution that avoids external critics at inference. Expect improved cross-benchmark transfer for iterative editing and reasoning tasks. Look elsewhere if you need a lightweight plug-in editor or a strict modular pipeline with separately maintainable verifier components: joint trajectory RL increases training complexity, requires rollout grouping and careful reward design, and may demand more compute and engineering to stabilize than conventional SFT or single-module RL baselines.