Addressing the systemic bottlenecks in multimodal document processing (MinerU, Docling) where variable-length child generation breaks GPU batching and lineage tracking, Peking University and OpenDCAI open-sourced RayOrch. Operating under a lineage-controlled multi-grain dataflow model with per-call FIFO ready queues, RayOrch delivers a 15.14x speedup scaling MinerU from 4 to 64 NVIDIA H20 GPUs, cutting end-to-end execution time by 13.1%~16.0% compared to Ray Data and 29.0% compared to Daft.

Key Takeaways

  • ✓End-to-End Lineage Preservation: Unlike legacy data systems that flatten records and lose lineage, RayOrch preserves hierarchical parent-child relationships and deterministic order throughout multimodal expansion pipelines.
  • ✓Cross-Parent FIFO GPU Batching: Introduces per-call FIFO ready queues that aggregate ready child tokens across disparate parents for optimal GPU saturation, paired with typed parent-scoped failure suppression.
  • ✓Scalability & Latency Reductions: On NVIDIA H20 clusters, scales MinerU across 4 to 64 GPUs with 15.14x speedup, cutting end-to-end latency by 13.1%~16.0% versus Ray Data and 29.0% versus Daft on Docling.
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点 / Background & Pain Points High-quality multimodal corpus preparation for foundation models requires expanding unstructured documents into dynamic, input-dependent child tokens (pages, tables, code blocks). Existing distributed data engines (Ray Data, Daft, Spark) either flatten hierarchies—forcing applications to manually manage lineage—or rely on coarse job partitions that fragment GPU batching and throttle cluster utilization. ### 架构亮点与底层机制 / Architectural Highlights Researchers from Peking University and OpenDCAI developed RayOrch, a lineage-controlled multi-grain dataflow programming model and runtime engine: 1. Declarative Expansions & Gathers: Validates ordered variable-cardinality expansions against corresponding gather sinks, tracking immutable ordinals and parent relations at the runtime kernel level; 2. Per-Call FIFO Ready Queues: Batches ready child items from disparate parents across GPU boundaries, retiring parents as soon as all dependent children terminate; 3. Typed Parent-Scoped Failure Suppression: Isolates localized errors by canceling undispatched siblings of a failed parent without halting unrelated parent branches; 4. Zero-Shuffle Lineage Reconstitution: Reconstructs outputs directly from structural ordinals rather than relying on batch boundary re-sorting. ### 权威 Benchmark 与实测跑分对比 / Benchmark & Evaluation Extensive benchmarks on NVIDIA H20 GPU clusters demonstrate clear systems advantages: - Linear GPU Scaling: Delivers 15.14x speedup scaling MinerU across 4 to 64 GPUs, and 7.82x on multimodal video pipelines across 8 to 64 GPUs; - End-to-End Latency: Cuts runtime by 13.1% vs Ray Data and 29.0% vs Daft on MinerU pipelines, and 16.0% vs Ray Data on Docling extraction; - Fault Recovery: Reduces wasted compute under dirty-record injection by over 80% through parent-scoped failure suppression. ### 开发者实战落地与开箱指南 / Developer Practical Guide - GitHub Repository: Accessible at https://github.com/OpenDCAI/RayOrch; - Runtime Integration: Operates as a distributed execution layer compatible with standard Ray clusters and container runners; - Adoption Advice: Wrap multi-stage document parsers using RayOrch's expansion and gather interfaces to maximize GPU memory bandwidth utilization during pre-training corpus preparation.