AIAny
Icon for item

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

Analyzes how on-policy distillation (OPD) transfers teacher LLM capabilities to student models across in-domain shifts, cross-domain transfer, and multi-teacher settings. Key findings: OPD conveys reasoning patterns rather than specific answers, same-origin teacher-student pairs generalize broadly, and multi-teacher combinations induce mixture-dependent tradeoffs.

Introduction

Why this matters OPD is widely used to transfer capabilities from stronger teacher LLMs to smaller students by supervising trajectories sampled from the student. Yet prior evaluations focus narrowly on single domains or benchmarks close to training data. This paper isolates one generalization factor at a time to reveal when OPD truly transfers reasoning behavior versus when it merely fits the trained distribution.

Key Findings
  • OPD transfers reasoning patterns, not just solutions: Training-problem difficulty has little effect — even problems the teacher never solves can help the student learn teacher-like reasoning. So what: supervision from OPD shapes the student’s internal inference style, reducing reliance on memorized answers.
  • Origin matters for breadth of transfer: Same-origin teacher-student pairs move the student close to the teacher across languages (English→Chinese), reasoning horizons (short→long), and other domains; cross-origin teachers mainly improve performance on the trained distribution. So what: model lineage and pretraining/architecture similarity can be more decisive than raw teacher performance when aiming for broad generalization.
  • Multi-teacher OPD is a double-edged sword: routing prompts to domain experts cannot fully confine each teacher’s influence, so combining experts produces a mixture-dependent seesaw among capabilities rather than independent composition. So what: naive expert routing can cause unexpected capability tradeoffs and requires diagnostic strategies.
Who it helps and tradeoffs

Great fit if you need to understand or deploy OPD for capability transfer across languages, reasoning horizons, or domains and want principled guidance on multi-teacher setups. Look elsewhere if you only care about short-term, within-distribution gains from offline distillation—OPD’s strengths lie in behavioral transfer rather than brute-force performance on a single benchmark.

Where it fits

The study complements OPD and model-distillation literature by emphasizing mechanism and practical diagnostics: it explains when OPD will generalize beyond training prompts and when teacher selection and origin considerations are more important than training-set difficulty.

Information

  • Websitearxiv.org
  • OrganizationsUniversity of Science and Technology of China, Peking University, IQuest Research, MBZUAI, Zhejiang University
  • AuthorsZhaoyi Li, Deyang Kong, Yuan Wei, Evan Yang, Ranran Shen, Mahardika Krisna Ihsani, Ming Yang, Wei Zhang, Chuan Hao, Jian Yang …
  • Published date2026/08/17

More Items

Proposes TRACE, an FP4 quantization framework for RL of MoE LLMs that uses rollout-side FP4 outcomes to guide training-side rounding and caches mantissa/scale from deeper layers to limit overhead. Preserves BF16-level RL performance while enabling up to 5.4× rollout speedup.

Analyzes cross-tokenizer on-policy distillation (OPD) and shows that focusing supervision on strictly aligned token positions—or a student-selected top-16 subset of the shared vocabulary—retains most distillation gains, while adding span-level MSE supervision reduces downstream accuracy.

Adapts LLM agents online by self-distilling verified execution trajectories into persistent LoRA weights during deployment to improve success and efficiency on long‑horizon tasks. Uses a frozen stable copy as a privileged teacher to predict hindsight next‑token distributions and filters invalid-action turns so experience consolidates without external solutions or memory retrieval.