AIAny
Icon for item

Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

Analyzes cross-tokenizer on-policy distillation (OPD) and shows that focusing supervision on strictly aligned token positions—or a student-selected top-16 subset of the shared vocabulary—retains most distillation gains, while adding span-level MSE supervision reduces downstream accuracy.

Introduction

Why this matters

Maximizing alignment coverage across different tokenizers feels like an obvious way to recover more supervision when distilling large models, but this paper shows that broader structural coverage can introduce weak or conflicting training signals that harm final accuracy. The core insight is that compact, reliable supervision at strictly aligned 1:1 token positions already captures most predictive mass and learning benefit for heterogeneous teacher–student pairs in math reasoning and code generation.

Key Findings
  • Strict 1:1 student-token coverage already spans most generated tokens despite static vocabulary mismatch; measured strict coverage is 85.57–96.98% across runs, while static vocabulary Jaccard overlap is only 39.49–64.87% — so token-level alignment matters more than raw vocab overlap.
  • The shared vocabulary at strict positions carries nearly all predictive probability (shared vocab retained 99.69–99.90% of teacher mass and 98.99–99.81% of student mass on average), implying excluded vocab entries rarely matter for distribution matching.
  • A compact student-selected top-k (k=16) subset at each strict position preserves at least 96% of the full-average improvement from full shared-vocabulary OPD and retains ≥93.5% teacher / ≥94.6% student mass, demonstrating practical, efficient supervision choices.
  • Adding span-level MSE supervision to cover mismatch groups (complete coverage) consistently lowers accuracy: all 18 tested positive-weight settings reduced full-average accuracy. Diagnostics show span gradients have weak or negative directional agreement with strict gradients and grow in magnitude during training, suggesting conflicting learning signals.
Who it's for + Tradeoffs

Great fit if you want a practical recipe for cross-tokenizer OPD that favors reliable, compact supervision over maximal coverage — especially for instruction-tuned or base student models in math reasoning and code generation tasks. The paper provides empirical numbers (coverage, retained mass, k=16 results) to guide design choices.

Look elsewhere if your priority is exhaustive structural coverage for theoretical completeness, or if your application demands span-level probability calibration despite possible accuracy regressions; span MSE can provide full structural supervision but at the cost of downstream performance.

Where It Fits

This work sits between cross-tokenizer distribution-matching methods (shared-vocab OPD, byte-prefix marginalization) and practical distillation engineering: it reframes the objective from "cover more" to "supervise reliably," offering actionable subset-selection and loss-weighting guidance for practitioners.

Information

  • Websitearxiv.org
  • OrganizationsGuohua Liu, Yuewei Zhang, Alibaba Cloud Computing
  • AuthorsBingxi Hou, Guochao Jiang, Guofeng Quan, Weiqing Li, Wenfeng Feng, Guohua Liu, Yuewei Zhang
  • Published date2026/10/06

More Items

Proposes TRACE, an FP4 quantization framework for RL of MoE LLMs that uses rollout-side FP4 outcomes to guide training-side rounding and caches mantissa/scale from deeper layers to limit overhead. Preserves BF16-level RL performance while enabling up to 5.4× rollout speedup.

Adaptive-granularity post-retrieval evidence compressor for multimodal RAG that selects variable-sized regions from retrieved text, tables, images, and videos to retain necessary context while cutting irrelevant content. Key features include a hierarchy-based node encoder with parent-relative refinement and a critic that issues targeted follow-up retrievals; yields higher QA accuracy across five benchmarks and reduces reader-input tokens by ~14–28%.

Adapts LLM agents online by self-distilling verified execution trajectories into persistent LoRA weights during deployment to improve success and efficiency on long‑horizon tasks. Uses a frozen stable copy as a privileged teacher to predict hindsight next‑token distributions and filters invalid-action turns so experience consolidates without external solutions or memory retrieval.