Extrapolates RL-induced representation residuals into student hidden states during on-policy distillation: at each layer and token, it regresses the student beyond the teacher along the teacher’s RL-induced direction to improve stability and empirical performance.
Uses a multimodal model's own critiques as privileged context and applies on-policy self-distillation over diffusion sampling trajectories to internalize corrective guidance, improving text-to-image generation without an external teacher; shows measurable gains on GenEval and GenEval2.