Most multi-teacher on-policy distillation (MOPD) work focuses on which specialist teaches each prompt, but neglects how strongly each specialist's token-level feedback should move the shared student. This paper shows that unequal feedback spread across domains (instruction-following feedback being much more diffuse than mathematics feedback) skews the optimization budget and prevents the student from inheriting some specialist gains. The core insight: deciding who teaches is insufficient—you must also decide how strongly each teacher's feedback counts.
Key Findings
- Measuring each domain's empirical feedback spread and rescaling teacher supervision by that spread (Domain-Normalized MOPD, DN-MOPD) reduces dominance from diffuse domains and restores balance across specialties.
- Across three model sizes (Qwen3.5 variants), six public benchmarks, three random seeds and two answer-length limits, DN-MOPD improves average performance over standard MOPD and recovers most of the mathematics specialist's lost advantage.
- Ablations with fixed domain weights indicate the primary benefit comes from attenuating instruction-following feedback rather than simply amplifying mathematics feedback; fixed weights close to DN-MOPD's measured scales perform similarly.
Method overview
- At each token, compute a per-domain measure of feedback spread (how widely the teacher's corrections are distributed) and use it to normalize that domain's contribution to the student's gradient updates.
- The approach keeps standard routing (which teacher is selected) but rescales the magnitude of the supervision so token-share and optimization budget are balanced across domains.
Who it's for and tradeoffs
Great fit if you train a single consolidated LLM from multiple RL-tuned specialists and observe capability imbalances (e.g., instruction-following dominating math). DN-MOPD is a low-friction, label-free adjustment to existing MOPD pipelines that preserves routing choices while rebalancing learning signals. Look elsewhere if imbalance stems from mislabeled data, fundamentally incompatible model capacity, or when per-domain sequence-length and convergence differences are already corrected by other scheduling mechanisms; DN-MOPD addresses feedback-scale imbalance specifically rather than all multi-domain training pathologies.