AIAny
Icon for item

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

Extrapolates RL-induced representation residuals into student hidden states during on-policy distillation: at each layer and token, it regresses the student beyond the teacher along the teacher’s RL-induced direction to improve stability and empirical performance.

Introduction

Why this matters Many on-policy distillation methods extrapolate an implicit reward in output (logit) space, but the LM head both attenuates and anisotropically filters representation changes and sampled log-ratio signals amplify noise. That makes output-space extrapolation unstable, especially when the RL teacher is close to its base checkpoint. The key insight of this paper is that RL produces a measurable directional shift in internal representations, and extrapolating that shift in representation space avoids the LM-head bottleneck and the high-variance token-probability noise.

Key Findings
  • RIDE (RL-Induced Direction Extrapolation) computes the per-layer, per-token residual between an RL-trained teacher and its pre-RL checkpoint, and regresses the student’s hidden states toward targets displaced beyond the teacher along that residual. This moves the student along the teacher’s learned direction while keeping it near the teacher.
  • The regression objective is equivalent (conditioned on a sampled trajectory) to maximizing a linear directional reward defined by the residual with a quadratic penalty centered on the teacher; this makes the optimization explicit about directionality and deviation control.
  • Across four base/teacher pairs spanning different scales, architectures, and pretraining lineages, RIDE approaches or exceeds the RL-trained teacher on every pair and is the only method whose mean performance does so; it consistently outperforms output-space extrapolation, which can degrade students when the teacher is close to its base checkpoint.
  • RIDE reduces sensitivity to noisy sampled-token log-probabilities and to the LM-head’s anisotropic mapping, improving stability when applying extrapolation-style objectives.
Who it’s for and tradeoffs

Great fit if you want to transfer RL gains from a teacher into a stronger student without running sparse-reward RL on the student, and you can access both the RL-trained teacher and its pre-RL checkpoint. RIDE is especially relevant when output-space extrapolation is unstable or when internal representation shifts are expected to carry the RL signal. Look elsewhere if you cannot access the teacher’s pre-RL checkpoint or you need a purely output-space method compatible with workflows that only expose logits; RIDE also requires incorporating representation-level losses into training, which changes memory/computation trade-offs compared to purely output-space objectives.

Information

  • Websitearxiv.org
  • AuthorsHao Li, MeiJia Chen, Weijie Ren, Donghan Li, Zijun Tian, Jingchun Huang, Naibo Wang
  • Published date2026/09/29

More Items

Trains a single unified multimodal model with reinforcement learning to perform end-to-end self-reflection and iterative image repair — jointly learning the diagnostic (textual) reflection and the flow-based image revisions so credit propagates across rounds without an external verifier.

Normalizes each domain's teacher feedback spread during multi-teacher on-policy distillation so that no domain (e.g., instruction following) overwhelms others, improving student recovery of specialist skills and raising average scores across benchmarks.

Introduces a rubric-based approach for video reward modeling that generates explicit, query-adaptive evaluation criteria before scoring to reduce scalar drift; proposes RGPO, a two-stage training (seed warm-up + joint optimization) that achieves strong pointwise and pairwise evaluation with high data efficiency.