Why this matters
Production coding and assistant agents create serving workloads that differ from single-request LLM traces: contexts grow large while outputs stay short, tool calls interleave tightly with model requests, and cached input dominates token volume and cost. This release gives infrastructure researchers realistic, replayable traces—timing, block-level prefix IDs, tool latencies and session structure—without exposing message text, enabling experiments that were previously constrained to synthetic benchmarks.
What Sets It Apart
- Real production scale and diversity: 16 weeks covering 12,002 reconstructed sessions across 14 harnesses, with per-request timing, provider-anonymized model IDs, and 1.2M tool calls—suitable for system-level evaluation rather than toy scenarios.
- Block-level prefix representation: inputs are tokenized into chained 16-token blocks with opaque block IDs, letting you replay KV-cache hits/evictions and measure prefix reuse without releasing source prompts.
- Tool and latency fidelity: each tool call records category, latency, result token counts and nested subagent sessions, so you can study tool-induced gaps, tool parallelism, and their impact on tail latency and batching.
- Two convenient formats: raw zstd JSONL (one session per line with nested structure) and flattened Parquet tables (sessions, requests, tool_calls) for fast analytics with DuckDB, pandas or Polars.
Who It's For and Tradeoffs
Great fit if you are building or evaluating serving infrastructure: KV-cache policies, tiered cache designs, prefetching, scheduling/batching strategies, and agent-aware session management. The trace supports realistic arrival patterns, mutation semantics, and tool delays that benchmark suites often omit.
Look elsewhere if your goal is content-level NLP research (no message text is released) or model training: the dataset intentionally omits readable prompts/responses and replaces rare sensitive tool arguments with typed placeholders. Also note block IDs are local to each week (make them globally unique when replaying multiple weeks) and sessions never cross week boundaries; these conventions matter when designing multi-week replay experiments.
Practical notes: the dataset is released under CC BY 4.0, tokenized with o200k_base into 16-token blocks, and includes helper scripts and a cache-simulation artifact to reproduce the paper's figures and run prefix-cache experiments.