Analyzes cross-tokenizer on-policy distillation (OPD) and shows that focusing supervision on strictly aligned token positions—or a student-selected top-16 subset of the shared vocabulary—retains most distillation gains, while adding span-level MSE supervision reduces downstream accuracy.
Proposes TRACE, an FP4 quantization framework for RL of MoE LLMs that uses rollout-side FP4 outcomes to guide training-side rounding and caches mantissa/scale from deeper layers to limit overhead. Preserves BF16-level RL performance while enabling up to 5.4× rollout speedup.