Rollout generation during RL fine-tuning of large language models demands substantial compute and memory; moving rollouts to aggressive FP4 quantization can greatly accelerate and shrink rollouts but often creates a train–rollout policy mismatch that destabilizes learning — a problem amplified in MoE models where routing magnifies numerical differences.
Key Findings
- TRACE directly aligns the quantized train and rollout execution paths by recording rollout-side FP4 quantization outcomes (activations and KV states) and using them to guide training-side rounding decisions. So what? This reduces local train–rollout discrepancy rather than merely minimizing each path's error against a high-precision reference.
- Efficient quantization-information caching: TRACE selectively retains mantissa and scale metadata from deeper layers to avoid prohibitive storage/communication overhead. So what? The scheme preserves most FP4 rollout throughput while keeping per-token cache costs small (example: ~7.5 KB cached info per token for Qwen3.5-35B when retaining 1-bit mantissa + amortized scale for latter 20 layers).
- Empirical outcomes: on multiple large MoE models and long-horizon reasoning/coding RL tasks, TRACE enables joint FP4 weight/activation and FP4 KV-cache rollout with RL performance comparable to BF16 rollout, achieves up to 5.4× decoding throughput versus BF16 rollout, and yields stronger final FP4 task performance than post-hoc FP4 quantization of BF16-trained policies. So what? Practitioners can obtain large rollout speedups without sacrificing final policy quality.
Who It's For and Tradeoffs
Great fit if you run reinforcement learning or post-training policy tuning on MoE LLMs and need to reduce rollout cost while retaining policy fidelity. TRACE suits teams willing to add rollout-side recording and a modest caching mechanism to their training loop. Look elsewhere if you cannot modify rollout/training pipelines to record or transmit quantization metadata, or if your deployment cannot accommodate the extra implementation complexity (even though TRACE aims to keep runtime overhead low).