Multimodal models that both generate and assess their outputs open the door to self-supervised improvement during post-training or test-time compute. UniEvo-VL exploits that generating is harder than verifying: it converts model self-critiques into privileged prompts and distills the corrective effect into the original-generation policy so the model improves from the vanilla prompt alone.
Key Findings
- Distillation-from-critique: Treats a single multimodal model as teacher and student under different contexts (teacher conditions on a critique-augmented prompt, student sees the original prompt) and minimizes per-state divergence across denoising diffusion trajectories. This transfers corrective behavior without requiring corrected images or scalar rewards.
- Empirical gains: Builds on Qwen-image variants and reports native GenEval direct-generation improvements (e.g., 0.747 → 0.808) and GenEval2 Soft-TIFA gains (e.g., 32.97 → 35.53). Using stronger external critics can further raise the achievable ceiling under the same framework.
- Complementary to reflection: The evolved generator still benefits from an additional reflection pass at inference time, indicating that internalized experience and runtime feedback are additive rather than redundant.
Who it's for and tradeoffs
Great fit if you want to improve a unified vision-language generator without collecting new labeled images or relying on a separate large teacher: the approach leverages the model's own understanding to supply dense supervision along sampling trajectories. Look elsewhere if your primary goal is guaranteed uniform gains across all sub-tasks (the paper reports mixed results, e.g., text-rendering may not uniformly improve) or if you cannot run additional on-policy training iterations and trajectory-level distillation due to compute constraints.
How it works (brief)
UniEvo-VL collects model self-critiques, rewrites prompts to include corrective cues, then runs on-policy self-distillation (OPSD) where teacher and student generate along the same student sampling trajectory but under different prompt contexts. The loss matches denoising distributions at each intermediate state, encouraging the student to internalize the teacher's corrective steps while preserving inference-time conditions (student only sees the original prompt). Experimental results validate the method on compositional generation and text-rendering benchmarks.