Most agent benchmarks report only final scores; this dataset reveals the full step-by-step runs so you can replay, audit, or analyze agent behavior across long workflows.
What Sets It Apart
- Step-level, replayable traces: 7,366 trajectories (Holo4-27B and Holo4-35B-A3B) with each step containing the model’s reasoning, the chosen actions or tool calls, tool results, and screenshots. This lets researchers reproduce exactly what the agent saw and why it acted.
- Compact, developer-oriented layout: a top-level index (data/index.json) with one summary row per trajectory and per-trajectory files (data/t/
<id>.json) including task, steps, verifier result and token usage; screenshots are stored as img/<id>/<step>.webp. The index rows include fields like model, benchmark, run, task, instruction, success, score, duration_s and steps. - Open provenance and licensing: the bundle mirrors the runs behind Holo4 benchmark scores and preserves upstream task licenses (Apache-2.0, MIT, CC BY 4.0 where applicable). Sensitive values are masked and screenshots containing PII are replaced with placeholders.
Data layout & reproducibility
- Each trajectory is a trace of a single agent session; steps grow prompts and include tool definitions and results as the agent interacted with GUIs, APIs or a code sandbox. Screenshots are provided (or placeholders when masked) so replay engines can render the same visual context.
- The dataset is structured for automated replay and analysis: index.json for filtering and statistics, per-trajectory JSON for stepwise replay, and image files referenced by steps. Typical uses include error analysis, verifier-driven evaluation, prompt-growth studies, and benchmarking agent architectures and toolchains.
Who Should Use It and Trade-offs
- Great fit if you evaluate or develop multimodal, tool-using agents (debugging failures, measuring hallucinations, profiling prompt growth, or building verifiers). Also useful for researchers studying long-horizon workflows and agent-tool interactions.
- Look elsewhere if you only need static labeled examples (this is trace-centric) or if you require private/proprietary app traces—sensitive fields are masked and a few tasks were omitted. Replaying large long trajectories requires infrastructure for many image-augmented requests and paying attention to context-size growth when simulating the original agent.
Where it fits
- Complementary to benchmark score reports: use the traces to explain why a model achieved (or failed) a given score, to compare decision patterns across model sizes, or to build new verifiers and evaluation metrics that operate on step-level behavior.