paper web signal

Nereus reports 27.7% step-latency cut, 7.27x over OpenRLHF

TL;DR

  • Nereus reports cutting average step latency by 27.7% by adapting tensor and pipeline parallelism during a run rather than restarting.
  • On 8B PPO, the runtime claims 2.14--7.27x throughput over OpenRLHF and 1.10--1.47x over Verl across diverse clusters.
  • In a 1,000-step run reaching 1,024 GPUs, six online layout transitions consumed 0.079% of total run time, per the paper.

Adapting tensor and pipeline parallelism mid-run, rather than restarting with a new layout, cut the average step latency of an 8B PPO job by 27.7%, according to a new preprint on arXiv.

The system, called Nereus, is described by authors Songlin Jiang, Tuo Shi, Sitong Zhang, Zeke Wang, Mario Di Francesco and Bo Zhao as 'a cost-aware runtime that adapts RL post-training jobs into efficient execution plans.' The problem they name: in RL post-training, 'resource availability, sequence length, memory pressure, and stage bottlenecks' all move during a run, and 'an execution plan that was initially suitable can then become slow or even infeasible over time.'

The bigger number is throughput. Against OpenRLHF, the paper claims end-to-end 8B PPO gains of 2.14--7.27x across diverse clusters. Against Verl, the range narrows to 1.10--1.47x.

Mechanically, Nereus represents the distributed state of each model-stage replica as what the authors call an 'Elastic Model Unit', with a 'global transition graph' ordering the GPU transfers when the plan changes. In a 1,000-step run reaching 1,024 GPUs, they report 'six transitions consume 0.079% of total run time'.

The trace behind the 27.7% figure is described only as 'built from real data'. No customer, cluster or model lineage is named, and the preprint has not been peer-reviewed.