huggingface.co web signal

Alibaba TRACE Matches BF16 MoE RL at 5.4x FP4 Rollout Speed

TL;DR

  • TRACE matches BF16 reinforcement-learning scores on four Qwen MoE policies (35B-A3B up to 2.4T-A95B) while running FP4 rollout up to 5.4x faster at 128K output.
  • On Qwen3.5-35B-A3B reasoning tasks the average score climbs from 68.8 under QUADS to 75.3, with HMMT25 jumping from 59.0 to 70.0.
  • The 5.4x figure is measured on 4 GB200 GPUs; a full rollout-side quantization record would run to roughly 51 TB per RL step before TRACE's one-bit mantissa cache.

An Alibaba team reports an FP4 rollout scheme that matches BF16-level reinforcement-learning scores across four Qwen Mixture-of-Experts models, and runs the rollout step up to 5.4x faster at 128K output length.

The paper, TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models, is the latest entry in a run of RL post-training work we have been tracking in our fine-tuning feed. Its trick is to let the FP4 rounding choices made during rollout generation guide the matching FP4 rounding decisions during the training update, rather than quantize the two execution paths independently. On reasoning tasks with Qwen3.5-35B-A3B, the authors write that "TRACE improves the average score from 68.8 to 75.3. The improvement is especially significant on HMMT25, where the score increases from 59.0 to 70.0."

On coding and long-horizon RL, the picture is parity with BF16: 33.0 vs 33.4 on DeepSWE for Qwen3.5-122B-A10B, 70.6 vs 68.8 on Terminal-Bench for Qwen3.8-Flash-Next, and 90.2 vs 90.3 on GDPval for Qwen3.8-2.4T-A95B.

The savings come with housekeeping. The paper notes that storing the full rollout-side quantization record for one RL step on Qwen3.5-35B-A3B would run to "approximately 51 TB of rollout-side quantization information per step"; TRACE's workaround is to keep only one-bit mantissa data for the latter 20 layers, about 7.5 KB per generated token. The throughput numbers are measured on 4 GB200 GPUs for rollout and 72 GB200 GPUs end-to-end, so the headline multiplier is a Blackwell result; the paper publishes no comparable numbers for older hardware.