Why this matters
Maximizing alignment coverage across different tokenizers feels like an obvious way to recover more supervision when distilling large models, but this paper shows that broader structural coverage can introduce weak or conflicting training signals that harm final accuracy. The core insight is that compact, reliable supervision at strictly aligned 1:1 token positions already captures most predictive mass and learning benefit for heterogeneous teacher–student pairs in math reasoning and code generation.
Key Findings
- Strict 1:1 student-token coverage already spans most generated tokens despite static vocabulary mismatch; measured strict coverage is 85.57–96.98% across runs, while static vocabulary Jaccard overlap is only 39.49–64.87% — so token-level alignment matters more than raw vocab overlap.
- The shared vocabulary at strict positions carries nearly all predictive probability (shared vocab retained 99.69–99.90% of teacher mass and 98.99–99.81% of student mass on average), implying excluded vocab entries rarely matter for distribution matching.
- A compact student-selected top-k (k=16) subset at each strict position preserves at least 96% of the full-average improvement from full shared-vocabulary OPD and retains ≥93.5% teacher / ≥94.6% student mass, demonstrating practical, efficient supervision choices.
- Adding span-level MSE supervision to cover mismatch groups (complete coverage) consistently lowers accuracy: all 18 tested positive-weight settings reduced full-average accuracy. Diagnostics show span gradients have weak or negative directional agreement with strict gradients and grow in magnitude during training, suggesting conflicting learning signals.
Who it's for + Tradeoffs
Great fit if you want a practical recipe for cross-tokenizer OPD that favors reliable, compact supervision over maximal coverage — especially for instruction-tuned or base student models in math reasoning and code generation tasks. The paper provides empirical numbers (coverage, retained mass, k=16 results) to guide design choices.
Look elsewhere if your priority is exhaustive structural coverage for theoretical completeness, or if your application demands span-level probability calibration despite possible accuracy regressions; span MSE can provide full structural supervision but at the cost of downstream performance.
Where It Fits
This work sits between cross-tokenizer distribution-matching methods (shared-vocab OPD, byte-prefix marginalization) and practical distillation engineering: it reframes the objective from "cover more" to "supervise reliably," offering actionable subset-selection and loss-weighting guidance for practitioners.