AIAny
Icon for item

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

Proposes TRACE, an FP4 quantization framework for RL of MoE LLMs that uses rollout-side FP4 outcomes to guide training-side rounding and caches mantissa/scale from deeper layers to limit overhead. Preserves BF16-level RL performance while enabling up to 5.4× rollout speedup.

Introduction

Rollout generation during RL fine-tuning of large language models demands substantial compute and memory; moving rollouts to aggressive FP4 quantization can greatly accelerate and shrink rollouts but often creates a train–rollout policy mismatch that destabilizes learning — a problem amplified in MoE models where routing magnifies numerical differences.

Key Findings
  • TRACE directly aligns the quantized train and rollout execution paths by recording rollout-side FP4 quantization outcomes (activations and KV states) and using them to guide training-side rounding decisions. So what? This reduces local train–rollout discrepancy rather than merely minimizing each path's error against a high-precision reference.
  • Efficient quantization-information caching: TRACE selectively retains mantissa and scale metadata from deeper layers to avoid prohibitive storage/communication overhead. So what? The scheme preserves most FP4 rollout throughput while keeping per-token cache costs small (example: ~7.5 KB cached info per token for Qwen3.5-35B when retaining 1-bit mantissa + amortized scale for latter 20 layers).
  • Empirical outcomes: on multiple large MoE models and long-horizon reasoning/coding RL tasks, TRACE enables joint FP4 weight/activation and FP4 KV-cache rollout with RL performance comparable to BF16 rollout, achieves up to 5.4× decoding throughput versus BF16 rollout, and yields stronger final FP4 task performance than post-hoc FP4 quantization of BF16-trained policies. So what? Practitioners can obtain large rollout speedups without sacrificing final policy quality.
Who It's For and Tradeoffs

Great fit if you run reinforcement learning or post-training policy tuning on MoE LLMs and need to reduce rollout cost while retaining policy fidelity. TRACE suits teams willing to add rollout-side recording and a modest caching mechanism to their training loop. Look elsewhere if you cannot modify rollout/training pipelines to record or transmit quantization metadata, or if your deployment cannot accommodate the extra implementation complexity (even though TRACE aims to keep runtime overhead low).

Information

  • Websitearxiv.org
  • OrganizationsAffiliation: Alibaba Token Hub, Alibaba Group, Affiliation: Ohio State University
  • AuthorsXin Wang, Hao Yu, Zhengyang Zhuge, Bochao Mao, Zheng Li, Junda Feng, Yuyan Luo, Yi Zhang, Yizhong Cao, Mi Zhang …
  • Published date2026/10/06

More Items

Analyzes cross-tokenizer on-policy distillation (OPD) and shows that focusing supervision on strictly aligned token positions—or a student-selected top-16 subset of the shared vocabulary—retains most distillation gains, while adding span-level MSE supervision reduces downstream accuracy.

Adaptive-granularity post-retrieval evidence compressor for multimodal RAG that selects variable-sized regions from retrieved text, tables, images, and videos to retain necessary context while cutting irrelevant content. Key features include a hierarchy-based node encoder with parent-relative refinement and a critic that issues targeted follow-up retrievals; yields higher QA accuracy across five benchmarks and reduces reader-input tokens by ~14–28%.

Adapts LLM agents online by self-distilling verified execution trajectories into persistent LoRA weights during deployment to improve success and efficiency on long‑horizon tasks. Uses a frozen stable copy as a privileged teacher to predict hindsight next‑token distributions and filters invalid-action turns so experience consolidates without external solutions or memory retrieval.