Why this matters
Agents spend most of their runtime making small, repeatable decisions—should I call a tool, which tool, and are the arguments valid. This dataset turns upstream agent transcripts into 180k typed, single-answer decision rows so you can train tiny, millisecond-scale decision models instead of running a full LLM for every step.
What Sets It Apart
- Typed, bounded decisions: every row encodes a question with an explicit criteria set (choice or bounded 0–1), which makes supervision and evaluation straightforward and deterministic — ideal for classifiers, encoders, or Jev-style typed heads.
- Leakage-safe splits and provenance: train/validation/test splits are group-safe and each row carries pinned upstream source, revision and license metadata, so evaluation reflects real generalization rather than split leakage.
- Practical tooling surface: flat Parquet with 35 columns (state_json, request_json, criteria_json, gold_label/score, provenance) fits DuckDB/Polars/pandas workflows and can be used immediately to train tool routers, argument checkers, or preference models.
Who It's For
Great fit if you need a small, high-quality labeled corpus to train or benchmark low-latency decision components (tool routers, when-to-call gates, argument validators) that sit in front of expensive LLM calls. Look elsewhere if you need full agent transcripts with raw model outputs, or human-curated preference labels — many labels are programmatic/automated and reflect upstream rules rather than human adjudication.