RayOrch: Open FM Data-Prep Framework Achieves 15.14x Strong-Scaling on 64 H20 GPUs

Found first: a primary source the press has not covered yet.

A 16-author team has published RayOrch, an open-source framework for foundation-model data preparation that achieves a 15.14x processing-time strong-scaling speedup when scaling from 4 to 64 NVIDIA H20 GPUs on the MinerU document pipeline. The paper, arXiv:2609.18703, describes a system that preserves parent-child lineage through variable-cardinality GPU batching, a structural problem that existing frameworks like Ray Data and Daft do not specifically address. Code is available at github.com/OpenDCAI/RayOrch.

What the source says

Document and video pipelines expand each input unit into an ordered, input-dependent sequence of children (pages per PDF, clips per video), and batching across these boundaries breaks parent-child ordering and completion tracking. RayOrch introduces lineage-controlled multi-grain dataflows that record each child's parent, ordinal, and status so GPU batches can span parent boundaries without losing structural integrity. On MinerU, end-to-end time was 13.1% lower than Ray Data and 29.0% lower than Daft across 64 H20 GPUs. The video pipeline scaled 7.82x from 8 to 64 GPUs, covering 104,952 clips from 27,091 videos. On Docling with 4 H20 GPUs, RayOrch ran 16.0% faster than Ray Data on 2,000 PDFs.

Why it matters

Data preparation for large models is typically the step that bottlenecks GPU utilization before training begins, and most teams build bespoke pipelines that do not scale across node counts. RayOrch provides a general abstraction for the expansion pattern common to document OCR, video segmentation, and similar preprocessing workloads. The 15.14x speedup across a 16-fold increase in GPUs (4 to 64) means scaling efficiency stays close to linear, which is the figure that matters for labs planning to throw more hardware at ingestion. Affiliations are not clearly stated in the paper's arXiv abstract; the GitHub organization is OpenDCAI.