Most existing LLM serving stacks assume all active requests are synchronized at each decoding step. Token-level routing breaks that assumption: requests may switch target models per token, creating step desynchronization, fragmented batches, and high implementation complexity for per-step routing. This paper argues that separating routing logic (from the request's perspective) and execution (from the model's perspective) unlocks large efficiency gains for routed inference.
Key Findings
- Route-send-receive programming model: express per-request token-level routing logic from the request's point of view; the runtime translates that into asynchronous execution across models, simplifying developer effort and avoiding invasive server changes.
- Subserver architecture with handoff-resume: each candidate LLM runs in its own subserver (scheduler, runner, private KV cache). Asynchronous subservers avoid forcing fast models to wait for slower ones, mitigating step desynchronization.
- Delayed-batching scheduler with analytic tuning: a delayed-batching policy that reduces average batch admission latency; its hyperparameters are derived from a throughput-based mathematical model to balance batching efficiency and latency.
- Empirical throughput gains: across multiple routing algorithms, workloads, and model pairs, the system attains 2.01–64.15× higher decoding throughput compared to existing single-LLM serving systems.
Who it's for and tradeoffs
Great fit if you run heterogeneous model ensembles or routing algorithms that select different LLMs per token and you need production-grade throughput improvements without rewriting serving internals. It benefits researchers prototyping token-level routing and engineering teams optimizing cost-quality tradeoffs in multi-model serving.
Look elsewhere if your deployment is single-model or single-GPU with trivial concurrency (no routing), since the subserver and scheduling machinery adds deployment complexity and requires tuning of scheduling hyperparameters. Multi-node setups also require network and resource orchestration for per-model subservers.
Where it fits
Positions itself between single-LLM servers (which assume step-level synchronization) and bespoke multi-model orchestration hacks. It acts as a drop-in external server (OpenAI-compatible API) while internally scaling per-model execution and scheduling.
How it works (brief)
Developers implement three small functions to describe routing for a single request: route(), send(), receive(). The runtime launches one subserver per candidate LLM; each subserver schedules requests using delayed batching, runs model decoding with a private KV-cache pool, and communicates peer-to-peer for handoffs. The system-level scheduler uses a throughput model to set delay windows to balance batch sizes and latency.
Overall insight: shifting complexity from monolithic synchronized serving to asynchronous, model-centric subservers + analytically tuned delayed batching yields large practical throughput improvements for token-level routed LLM inference.