paper web signal

RayOrch scales MinerU 15.14x from 4 to 64 H20 GPUs

TL;DR

  • RayOrch reports a 15.14x speedup scaling MinerU from 4 to 64 NVIDIA H20 GPUs, and 7.82x for a video pipeline from 8 to 64 GPUs.
  • On MinerU, the paper claims end-to-end times 13.1% faster than Ray Data and 29.0% faster than Daft; on Docling, 16.0% faster than Ray Data.
  • The engine preserves parent-child lineage, ordinals, and terminal states so batched children can be regrouped without depending on completion order.

The paper reports scaling the MinerU document pipeline by 15.14x when moving from 4 to 64 NVIDIA H20 GPUs, a nearly linear result. On a video pipeline from 8 to 64 GPUs, the reported speedup is 7.82x.

The system, called RayOrch, targets a specific pain in foundation-model data prep: documents and videos "expand each parent item into an ordered and input dependent sequence of children, whose counts may be long tailed," per the abstract. Batching those children across parents while keeping track of which parent they belong to, and in what order, is the part existing engines handle awkwardly.

The authors write that RayOrch "preserves parent child relations throughout execution," using what they call "Per Call FIFO Ready Queues" to batch ready children across parents, and reconstructing results "from declared membership and ordinals rather than batch boundaries or completion order." Failures are typed and parent-scoped, which the abstract says "suppress undispatched siblings of the failed parent while allowing unrelated parents to continue."

Head-to-head, the abstract claims 13.1% faster end-to-end time than Ray Data and 29.0% faster than Daft on MinerU, and 16.0% faster than Ray Data on Docling. All numbers are self-reported on H20 GPUs. The abstract reports no runs on other GPU generations. Code is on GitHub at OpenDCAI/RayOrch.