AIAny
Icon for item

Holo4 Trajectories

Provides 7,366 recorded agent trajectories from H Company’s Holo4 benchmark runs, with step-level reasoning, actions, tool results, token usage and screenshots for replay and analysis. Bundled as JSON and image files for per-trajectory inspection and automated replay; released under Apache 2.0.

Introduction

Most agent benchmarks report only final scores; this dataset reveals the full step-by-step runs so you can replay, audit, or analyze agent behavior across long workflows.

What Sets It Apart
  • Step-level, replayable traces: 7,366 trajectories (Holo4-27B and Holo4-35B-A3B) with each step containing the model’s reasoning, the chosen actions or tool calls, tool results, and screenshots. This lets researchers reproduce exactly what the agent saw and why it acted.
  • Compact, developer-oriented layout: a top-level index (data/index.json) with one summary row per trajectory and per-trajectory files (data/t/<id>.json) including task, steps, verifier result and token usage; screenshots are stored as img/<id>/<step>.webp. The index rows include fields like model, benchmark, run, task, instruction, success, score, duration_s and steps.
  • Open provenance and licensing: the bundle mirrors the runs behind Holo4 benchmark scores and preserves upstream task licenses (Apache-2.0, MIT, CC BY 4.0 where applicable). Sensitive values are masked and screenshots containing PII are replaced with placeholders.
Data layout & reproducibility
  • Each trajectory is a trace of a single agent session; steps grow prompts and include tool definitions and results as the agent interacted with GUIs, APIs or a code sandbox. Screenshots are provided (or placeholders when masked) so replay engines can render the same visual context.
  • The dataset is structured for automated replay and analysis: index.json for filtering and statistics, per-trajectory JSON for stepwise replay, and image files referenced by steps. Typical uses include error analysis, verifier-driven evaluation, prompt-growth studies, and benchmarking agent architectures and toolchains.
Who Should Use It and Trade-offs
  • Great fit if you evaluate or develop multimodal, tool-using agents (debugging failures, measuring hallucinations, profiling prompt growth, or building verifiers). Also useful for researchers studying long-horizon workflows and agent-tool interactions.
  • Look elsewhere if you only need static labeled examples (this is trace-centric) or if you require private/proprietary app traces—sensitive fields are masked and a few tasks were omitted. Replaying large long trajectories requires infrastructure for many image-augmented requests and paying attention to context-size growth when simulating the original agent.
Where it fits
  • Complementary to benchmark score reports: use the traces to explain why a model achieved (or failed) a given score, to compare decision patterns across model sizes, or to build new verifiers and evaluation metrics that operate on step-level behavior.

Information

Categories

More Items

Hugging Face

Provides 150,000 source‑grounded decision examples for training models that pick options, judge yes/no propositions, or assign ordered scores. Each row pairs a 'state' with JEV-style typed questions (CHOICE/NOUL/SCORE); multiple configs and train/test splits are included, license mixed/unknown.

Hugging Face

Provides 23,625 semi-structured smart-contract audit findings (title, description, PoC, recommendation, normalized severity) for defensive-security research; requires cleaning, deduplication, and PoC filtering before model training.

Hugging Face

Provides over 1.1M hours of high-bandwidth, multichannel multilingual speech with segment- and word-level timestamps, English translations, and per-file metadata for ASR, TTS and audio-representation research. Preserves original 48kHz multichannel OPUS audio and is released under CC BY 3.0.