AIAny
Icon for item

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

Uses a multimodal model's own critiques as privileged context and applies on-policy self-distillation over diffusion sampling trajectories to internalize corrective guidance, improving text-to-image generation without an external teacher; shows measurable gains on GenEval and GenEval2.

Introduction

Multimodal models that both generate and assess their outputs open the door to self-supervised improvement during post-training or test-time compute. UniEvo-VL exploits that generating is harder than verifying: it converts model self-critiques into privileged prompts and distills the corrective effect into the original-generation policy so the model improves from the vanilla prompt alone.

Key Findings
  • Distillation-from-critique: Treats a single multimodal model as teacher and student under different contexts (teacher conditions on a critique-augmented prompt, student sees the original prompt) and minimizes per-state divergence across denoising diffusion trajectories. This transfers corrective behavior without requiring corrected images or scalar rewards.
  • Empirical gains: Builds on Qwen-image variants and reports native GenEval direct-generation improvements (e.g., 0.747 → 0.808) and GenEval2 Soft-TIFA gains (e.g., 32.97 → 35.53). Using stronger external critics can further raise the achievable ceiling under the same framework.
  • Complementary to reflection: The evolved generator still benefits from an additional reflection pass at inference time, indicating that internalized experience and runtime feedback are additive rather than redundant.
Who it's for and tradeoffs

Great fit if you want to improve a unified vision-language generator without collecting new labeled images or relying on a separate large teacher: the approach leverages the model's own understanding to supply dense supervision along sampling trajectories. Look elsewhere if your primary goal is guaranteed uniform gains across all sub-tasks (the paper reports mixed results, e.g., text-rendering may not uniformly improve) or if you cannot run additional on-policy training iterations and trajectory-level distillation due to compute constraints.

How it works (brief)

UniEvo-VL collects model self-critiques, rewrites prompts to include corrective cues, then runs on-policy self-distillation (OPSD) where teacher and student generate along the same student sampling trajectory but under different prompt contexts. The loss matches denoising distributions at each intermediate state, encouraging the student to internalize the teacher's corrective steps while preserving inference-time conditions (student only sees the original prompt). Experimental results validate the method on compositional generation and text-rendering benchmarks.

Information

  • Websitearxiv.org
  • OrganizationsStanford University, Johns Hopkins University, Independent Researcher, University of Toronto, University of Oxford, UC, Riverside, MatrAIx, University of Notre Dame, Carnegie Mellon University, UNC–Chapel Hill
  • AuthorsFang Wu, Da Xing, Yanjie Huang, Junxi Wang, Ji Wang, Hejia Geng, Guancheng Wan, Bowen Zuo, Xiaomin Li, Shixiang Tang …
  • Published date2026/09/30

More Items

Analyzes how proposer–solver loops in self-evolving search agents can develop shared errors (co-cheating) that inflate internal rewards; introduces Multi-Sample Verification and CrossFit (cross-fitted scoring with partitioned sources) to reduce false agreement and improve downstream search performance.

Groups visually grounded appearances of the same physical instance into persistent, retrievable “biographies” so agents can follow objects across hours or days for long-video question answering. Links identity-aware observations to episodic context and visual evidence; improves EgoLifeQA to 72.0% and increases evidence-window reach from 37.6% to 58.9%.

Evaluates multimodal memory in large audio language models by testing four acoustic evidence types and four memory operations across multi-session spoken histories, stratified across 8K–64K token context budgets.