Why this matters Many on-policy distillation methods extrapolate an implicit reward in output (logit) space, but the LM head both attenuates and anisotropically filters representation changes and sampled log-ratio signals amplify noise. That makes output-space extrapolation unstable, especially when the RL teacher is close to its base checkpoint. The key insight of this paper is that RL produces a measurable directional shift in internal representations, and extrapolating that shift in representation space avoids the LM-head bottleneck and the high-variance token-probability noise.
Key Findings
- RIDE (RL-Induced Direction Extrapolation) computes the per-layer, per-token residual between an RL-trained teacher and its pre-RL checkpoint, and regresses the student’s hidden states toward targets displaced beyond the teacher along that residual. This moves the student along the teacher’s learned direction while keeping it near the teacher.
- The regression objective is equivalent (conditioned on a sampled trajectory) to maximizing a linear directional reward defined by the residual with a quadratic penalty centered on the teacher; this makes the optimization explicit about directionality and deviation control.
- Across four base/teacher pairs spanning different scales, architectures, and pretraining lineages, RIDE approaches or exceeds the RL-trained teacher on every pair and is the only method whose mean performance does so; it consistently outperforms output-space extrapolation, which can degrade students when the teacher is close to its base checkpoint.
- RIDE reduces sensitivity to noisy sampled-token log-probabilities and to the LM-head’s anisotropic mapping, improving stability when applying extrapolation-style objectives.
Who it’s for and tradeoffs
Great fit if you want to transfer RL gains from a teacher into a stronger student without running sparse-reward RL on the student, and you can access both the RL-trained teacher and its pre-RL checkpoint. RIDE is especially relevant when output-space extrapolation is unstable or when internal representation shifts are expected to carry the RL signal. Look elsewhere if you cannot access the teacher’s pre-RL checkpoint or you need a purely output-space method compatible with workflows that only expose logits; RIDE also requires incorporating representation-level losses into training, which changes memory/computation trade-offs compared to purely output-space objectives.