AIAny
Icon for item

On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

Investigates how rollout policy, token-level KL direction, and learning rate each affect LLM distillation across Llama3 and Qwen2.5 on reasoning tasks; finds KL direction and learning rate dominate outcomes while rollout policy has a modest effect.

Introduction

Why this matters and the core insight

Many practitioners assume on-policy rollouts are inherently beneficial for distillation, protecting prior capabilities and producing sparser updates. This paper shows a different, more nuanced picture: when rollout policy, token-level KL direction, and learning rate are varied independently, KL direction and learning rate drive most differences in accuracy, forgetting, and update sparsity, while rollout policy has only a modest role.

Key Findings
  • Forward vs reverse token-level KL: Forward KL is robust to the choice of rollout policy and yields stable, strong performance; reverse KL is highly sensitive and tends to prefer student-generated (on-policy) rollouts — so the KL objective itself largely shapes output coverage and task accuracy.
  • Learning rate governs forgetting and update sparsity: Lower learning rates reduce catastrophic forgetting and produce much sparser parameter updates; differences from rollout policy are comparatively small.
  • No consistent on-policy advantage in final in-distribution accuracy: Across tasks and model families, on-policy rollouts do not reliably improve target-task accuracy, forgetting, or sparsity; any on-policy gains (e.g., on harder Countdown variants) may not persist after further RL-based training.
  • Robustness of conclusions: Results hold across Llama3 and Qwen2.5 families, ablations (no gradient clipping, sampled KL estimators), and tasks requiring longer reasoning chains — indicating generality beyond narrow setups.
Who it's for and when to avoid

Great fit if you design or evaluate LLM post-training pipelines and need to disentangle how objective choice and optimization settings affect distillation outcomes. The paper helps prioritize hyperparameters (KL direction, learning rate) over defaulting to on-policy data. Look elsewhere if your primary concern is deployment engineering, latency, or non-LLM domains — the study focuses on controlled LLM distillation experiments rather than system-level integration.

Methodological notes

Controlled strong-to-weak distillation experiments independently vary rollout policy, token-level KL direction, and learning rate (two rates tested) using Llama-3.1→3.2 and Qwen2.5 teacher→student pairs on scientific, medical, and arithmetic reasoning tasks. Training used short full-parameter fine-tuning runs with multiple seeds and targeted metrics for in-distribution accuracy, OOD forgetting, and parameter-update sparsity.

Information

  • Websitearxiv.org
  • OrganizationsAffiliation: University of Cambridge, Affiliation: Cambridge, UK
  • AuthorsJulianna Piskorz, Antonin Berthon, Mihaela van der Schaar
  • Published date2026/09/28

More Items

Calibrates multi-reward reinforcement learning by adaptively upweighting infrequently active rewards per rollout batch, so sparse objectives provide stronger signals when they matter. Proposes an inverse-square-root density correction and shows faster learning on tool-calling and math-reasoning tasks.

Models sequence generation by unmasking multiple tokens per denoising step and replaces a factorized reverse process with a mixture over discrete routing-based latents from an MoE backbone; improves few-step sampling quality without increasing active parameters.

Demonstrates that pretrained transformers typically use only ~1–3 lines of depth to follow reference chains, and that a task‑trained rank‑8 LoRA applied at one early layer (with all other weights frozen) can extend reference‑following to dozens or hundreds of lines while adding only a few ten‑thousand parameters.