Most foundation-model data prep pipelines expand each input (document, video) into an ordered, input-dependent sequence of children whose counts are long-tailed. The core problem is not raw parallelism but preserving parent–child lineage, child order, and deterministic result routing while still enabling GPUs to form efficient cross-input batches. RayOrch's key insight is to make lineage and ordinal metadata first-class: the compiler validates declared expansions and gathers, and the runtime records child membership, immediate parents, immutable ordinals, and terminal states so batching, completion, and reassembly can be decoupled from physical execution order.
Key Findings
- Lineage-first execution: Declared expansions + per-call FIFO ready queues let the runtime batch ready children from many parents without losing ownership or order. So what: GPUs get large microbatches for throughput while per-parent correctness and ordering are preserved.
- Completion-driven assembly: Gathers reconstruct results by membership and ordinal rather than batch boundaries or completion order. So what: a parent can advance as soon as all its required children are terminal, avoiding head-of-line blocking from slower sibling inputs.
- Parent-scoped failure semantics: Typed failures suppress undispatched siblings of a failed parent while letting unrelated parents continue. So what: failures are contained, improving pipeline robustness for large heterogeneous corpora.
- Empirical benefits: On NVIDIA H20 GPUs RayOrch reports 15.14× speedup scaling MinerU from 4→64 GPUs and 7.82× scaling a video pipeline from 8→64 GPUs; end-to-end time reductions include 13.1% vs Ray Data and 29.0% vs Daft on MinerU, and 16.0% vs Ray Data on Docling. So what: for long-tailed, multi-grain data-prep workloads, RayOrch materially improves GPU utilization and overall turnaround.
Who it's for and trade-offs
Great fit if: you build large-scale data-preparation pipelines for foundation models where inputs expand into ordered variable-size child tasks (e.g., PDF pages, video frames), need deterministic reassembly per parent, and want to maximize cross-input GPU batching without shifting lineage bookkeeping into application code.
Look elsewhere if: your workload is simple, embarrassingly parallel map-only processing with no ordering or parent-scoped aggregation needs, or you cannot adopt Ray-based runtime semantics. RayOrch adds runtime bookkeeping and relies on Ray actor pools and per-stage resource configuration, so it is most beneficial when the added orchestration complexity pays off in batching and utilization gains.
Where it fits
Positioned between coarse-grained job systems (which hide parallelism) and flat record APIs (which force application-level regrouping). RayOrch targets pipeline-parallel, multi-model, multimodal workloads (PDF understanding, video processing, multi-model vision stacks, multi-stage LLM inference) that follow a 1→M→1 pattern and benefit from lineage-aware batching and ordered gathers.