AIAny
MLOps2026
Icon for item

RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation

Provides a programming model and distributed runtime that preserves parent–child lineage and ordered variable-cardinality expansions for 1→M→1 dataflow pipelines, enabling completion-driven cross-input GPU batching and ordered gathers for foundation-model data preparation. Demonstrates multi-GPU scaling speedups and lower end-to-end time versus Ray Data and Daft.

Introduction

Most foundation-model data prep pipelines expand each input (document, video) into an ordered, input-dependent sequence of children whose counts are long-tailed. The core problem is not raw parallelism but preserving parent–child lineage, child order, and deterministic result routing while still enabling GPUs to form efficient cross-input batches. RayOrch's key insight is to make lineage and ordinal metadata first-class: the compiler validates declared expansions and gathers, and the runtime records child membership, immediate parents, immutable ordinals, and terminal states so batching, completion, and reassembly can be decoupled from physical execution order.

Key Findings
  • Lineage-first execution: Declared expansions + per-call FIFO ready queues let the runtime batch ready children from many parents without losing ownership or order. So what: GPUs get large microbatches for throughput while per-parent correctness and ordering are preserved.
  • Completion-driven assembly: Gathers reconstruct results by membership and ordinal rather than batch boundaries or completion order. So what: a parent can advance as soon as all its required children are terminal, avoiding head-of-line blocking from slower sibling inputs.
  • Parent-scoped failure semantics: Typed failures suppress undispatched siblings of a failed parent while letting unrelated parents continue. So what: failures are contained, improving pipeline robustness for large heterogeneous corpora.
  • Empirical benefits: On NVIDIA H20 GPUs RayOrch reports 15.14× speedup scaling MinerU from 4→64 GPUs and 7.82× scaling a video pipeline from 8→64 GPUs; end-to-end time reductions include 13.1% vs Ray Data and 29.0% vs Daft on MinerU, and 16.0% vs Ray Data on Docling. So what: for long-tailed, multi-grain data-prep workloads, RayOrch materially improves GPU utilization and overall turnaround.
Who it's for and trade-offs

Great fit if: you build large-scale data-preparation pipelines for foundation models where inputs expand into ordered variable-size child tasks (e.g., PDF pages, video frames), need deterministic reassembly per parent, and want to maximize cross-input GPU batching without shifting lineage bookkeeping into application code.

Look elsewhere if: your workload is simple, embarrassingly parallel map-only processing with no ordering or parent-scoped aggregation needs, or you cannot adopt Ray-based runtime semantics. RayOrch adds runtime bookkeeping and relies on Ray actor pools and per-stage resource configuration, so it is most beneficial when the added orchestration complexity pays off in batching and utilization gains.

Where it fits

Positioned between coarse-grained job systems (which hide parallelism) and flat record APIs (which force application-level regrouping). RayOrch targets pipeline-parallel, multi-model, multimodal workloads (PDF understanding, video processing, multi-model vision stacks, multi-stage LLM inference) that follow a 1→M→1 pattern and benefit from lineage-aware batching and ordered gathers.

Information

  • Websitearxiv.org
  • AuthorsXiaochen Ma, Zimo Meng, Junzhu Liang, Youhe Jiang, Yue Cheng, Hao Liang, Bohan Zeng, Dengchun Li, Lu Ma, Zhengyang Zhao …
  • Published date2026/09/16

Categories

More Items

GitHub
AI Infra2026

Provides an end-to-end platform to evaluate, observe, protect, and optimize LLM and AI agent deployments. Integrates OpenTelemetry tracing, 50+ evaluation metrics, agent simulations, an OpenAI‑compatible gateway, and guardrails; self‑hostable under Apache 2.0.

GitHub
AI Train2026

Provides a one-command CLI to fine-tune and post-train LLMs, with layer streaming that lets an 8B model be fine-tuned on a 4 GB laptop GPU. Auto-configures quantization, LoRA adapters, batching and evaluation gates, and supports export and serving workflows.

GitHub
AI Infra2023

Curated learning hub that aggregates roadmaps, tutorials, bootcamps, books, projects, and tool recommendations for learning data engineering and production data infrastructure. Focuses on practical applied learning (projects, interview prep, community links) rather than code libraries.