AIAny
Icon for item

TokenRouter: Efficient Serving System for Token-Level LLM Routing

Serves token-level routed LLM inference by dispatching requests to per-model asynchronous subservers and using delayed-batching scheduling to reduce admission latency and step desynchronization. Exposes a request-centric route-send-receive API and reports 2.01–64.15× decoding throughput gains versus single-LLM servers.

Introduction

Most existing LLM serving stacks assume all active requests are synchronized at each decoding step. Token-level routing breaks that assumption: requests may switch target models per token, creating step desynchronization, fragmented batches, and high implementation complexity for per-step routing. This paper argues that separating routing logic (from the request's perspective) and execution (from the model's perspective) unlocks large efficiency gains for routed inference.

Key Findings
  • Route-send-receive programming model: express per-request token-level routing logic from the request's point of view; the runtime translates that into asynchronous execution across models, simplifying developer effort and avoiding invasive server changes.
  • Subserver architecture with handoff-resume: each candidate LLM runs in its own subserver (scheduler, runner, private KV cache). Asynchronous subservers avoid forcing fast models to wait for slower ones, mitigating step desynchronization.
  • Delayed-batching scheduler with analytic tuning: a delayed-batching policy that reduces average batch admission latency; its hyperparameters are derived from a throughput-based mathematical model to balance batching efficiency and latency.
  • Empirical throughput gains: across multiple routing algorithms, workloads, and model pairs, the system attains 2.01–64.15× higher decoding throughput compared to existing single-LLM serving systems.
Who it's for and tradeoffs

Great fit if you run heterogeneous model ensembles or routing algorithms that select different LLMs per token and you need production-grade throughput improvements without rewriting serving internals. It benefits researchers prototyping token-level routing and engineering teams optimizing cost-quality tradeoffs in multi-model serving.

Look elsewhere if your deployment is single-model or single-GPU with trivial concurrency (no routing), since the subserver and scheduling machinery adds deployment complexity and requires tuning of scheduling hyperparameters. Multi-node setups also require network and resource orchestration for per-model subservers.

Where it fits

Positions itself between single-LLM servers (which assume step-level synchronization) and bespoke multi-model orchestration hacks. It acts as a drop-in external server (OpenAI-compatible API) while internally scaling per-model execution and scheduling.

How it works (brief)

Developers implement three small functions to describe routing for a single request: route(), send(), receive(). The runtime launches one subserver per candidate LLM; each subserver schedules requests using delayed batching, runs model decoding with a private KV-cache pool, and communicates peer-to-peer for handoffs. The system-level scheduler uses a throughput model to set delay windows to balance batch sizes and latency.

Overall insight: shifting complexity from monolithic synchronized serving to asynchronous, model-centric subservers + analytically tuned delayed batching yields large practical throughput improvements for token-level routed LLM inference.

Information

  • Websitearxiv.org
  • OrganizationsTsinghua University, Carnegie Mellon University
  • AuthorsTianyu Fu, Tengxuan Liu, Ruoxi Wang, Yixin Dong, Yi Ge, Yichen You, Yu Wang
  • Published date2026/10/08

More Items

Converts historical interaction traces into a reusable, queryable “worldbook” and runs a language-based world model agent (Trace2Env) as the environment for LLM agents — enabling stateful, grounded simulation with improved next-observation fidelity and long-horizon consistency.

Analyzes why self-evolving reasoning models collapse under repeated self-training and proposes R-Quest: a feedback-driven pipeline that trains solvers to reject invalid questions and uses a frozen base model to detect task-level repetition, filtering training data to sustain multi-round gains.

Compresses recurrent states in linear-attention LLMs via spatial–temporal post-training quantization, allocating bits by error lifetime and per-row impact to preserve accuracy (6-bit ≈ FP32) while cutting serving memory up to 68.7%.