Why this matters and the core insight
Many practitioners assume on-policy rollouts are inherently beneficial for distillation, protecting prior capabilities and producing sparser updates. This paper shows a different, more nuanced picture: when rollout policy, token-level KL direction, and learning rate are varied independently, KL direction and learning rate drive most differences in accuracy, forgetting, and update sparsity, while rollout policy has only a modest role.
Key Findings
- Forward vs reverse token-level KL: Forward KL is robust to the choice of rollout policy and yields stable, strong performance; reverse KL is highly sensitive and tends to prefer student-generated (on-policy) rollouts — so the KL objective itself largely shapes output coverage and task accuracy.
- Learning rate governs forgetting and update sparsity: Lower learning rates reduce catastrophic forgetting and produce much sparser parameter updates; differences from rollout policy are comparatively small.
- No consistent on-policy advantage in final in-distribution accuracy: Across tasks and model families, on-policy rollouts do not reliably improve target-task accuracy, forgetting, or sparsity; any on-policy gains (e.g., on harder Countdown variants) may not persist after further RL-based training.
- Robustness of conclusions: Results hold across Llama3 and Qwen2.5 families, ablations (no gradient clipping, sampled KL estimators), and tasks requiring longer reasoning chains — indicating generality beyond narrow setups.
Who it's for and when to avoid
Great fit if you design or evaluate LLM post-training pipelines and need to disentangle how objective choice and optimization settings affect distillation outcomes. The paper helps prioritize hyperparameters (KL direction, learning rate) over defaulting to on-policy data. Look elsewhere if your primary concern is deployment engineering, latency, or non-LLM domains — the study focuses on controlled LLM distillation experiments rather than system-level integration.
Methodological notes
Controlled strong-to-weak distillation experiments independently vary rollout policy, token-level KL direction, and learning rate (two rates tested) using Llama-3.1→3.2 and Qwen2.5 teacher→student pairs on scientific, medical, and arithmetic reasoning tasks. Training used short full-parameter fine-tuning runs with multiple seeds and targeted metrics for in-distribution accuracy, OOD forgetting, and parameter-update sparsity.