Provides 16 weeks of anonymized production agent-session traces (12,002 sessions, ~1.19M LLM requests, ~1.21M tool calls, 209B input tokens) released as block-level prefix IDs plus flattened Parquet tables for KV-cache, scheduling and serving-system research.
Serves token-level routed LLM inference by dispatching requests to per-model asynchronous subservers and using delayed-batching scheduling to reduce admission latency and step desynchronization. Exposes a request-centric route-send-receive API and reports 2.01–64.15× decoding throughput gains versus single-LLM servers.