Nereus reports 27.7% step-latency cut, 7.27x over OpenRLHF
TL;DR
- Nereus reports cutting average step latency by 27.7% by adapting tensor and pipeline parallelism during a run rather than restarting.
- On 8B PPO, the runtime claims 2.14--7.27x throughput over OpenRLHF and 1.10--1.47x over Verl across diverse clusters.
- In a 1,000-step run reaching 1,024 GPUs, six online layout transitions consumed 0.079% of total run time, per the paper.
Adapting tensor and pipeline parallelism mid-run, rather than restarting with a new layout, cut the average step latency of an 8B PPO job by 27.7%, according to a new preprint on arXiv.
The system, called Nereus, is described by authors Songlin Jiang, Tuo Shi, Sitong Zhang, Zeke Wang, Mario Di Francesco and Bo Zhao as 'a cost-aware runtime that adapts RL post-training jobs into efficient execution plans.' The problem they name: in RL post-training, 'resource availability, sequence length, memory pressure, and stage bottlenecks' all move during a run, and 'an execution plan that was initially suitable can then become slow or even infeasible over time.'
The bigger number is throughput. Against OpenRLHF, the paper claims end-to-end 8B PPO gains of 2.14--7.27x across diverse clusters. Against Verl, the range narrows to 1.10--1.47x.
Mechanically, Nereus represents the distributed state of each model-stage replica as what the authors call an 'Elastic Model Unit', with a 'global transition graph' ordering the GPU transfers when the plan changes. In a 1,000-step run reaching 1,024 GPUs, they report 'six transitions consume 0.079% of total run time'.
The trace behind the 27.7% figure is described only as 'built from real data'. No customer, cluster or model lineage is named, and the preprint has not been peer-reviewed.
Originally reported by paper
Read the original article →Original headline: Nereus Cuts RL Post-Training Step Latency 27.7%, Runs Up to 7x Faster Than OpenRLHF on 8B PPO