Extrapolates RL-induced representation residuals into student hidden states during on-policy distillation: at each layer and token, it regresses the student beyond the teacher along the teacher’s RL-induced direction to improve stability and empirical performance.