AIAny
Icon for item

Fable 5 Boeing 747 - Claude Code session trace

Contains a sanitized Claude Code (Fable 5) JSONL transcript of a session that procedurally built a Boeing 747 in Three.js, including assistant messages, tool calls, and base64 screenshots — useful for studying agent trace, tool use, and vision self‑verification workflows.

Introduction

This dataset exposes a compact, real-world example of a multimodal agent execution: a Claude Code (Fable 5) session that iteratively built a procedural Boeing 747 in Three.js using a vision‑enabled self‑verification loop.

What Sets It Apart
  • Full-session JSONL transcript including user prompts, assistant thinking, tool calls, tool outputs, and embedded base64 screenshots — not just high‑level logs. This makes it possible to replay or analyze the agent’s stepwise decisions and visual checks.
  • Focused multimodal workflow: the agent uses rendered screenshots as part of an autonomous verify‑and‑iterate loop, illustrating practical patterns for vision‑guided code generation and self‑validation.
  • Small, sanitized, MIT‑licensed dataset formatted for easy inspection with standard tools (pandas/polars), ideal as a concrete example rather than a large corpus.
Who It's For and Tradeoffs

Great fit if you study AI agent behavior, tool use, or multimodal self‑verification patterns — for researchers, educators, or engineers prototyping agent debugging workflows. Look elsewhere if you need large-scale labelled datasets, broad agent benchmarks, or raw production logs: this is a single, focused session (small size) and contains redactions/placeholders for private paths and environment details.

Where It Fits

Compliments larger agent‑trace collections by providing a clear, executable example of visual inspection integrated into an autonomous generation loop. Use it to prototype replay tools, visualization dashboards, or case studies about agent verification strategies.

More Items

Evaluates multimodal context learning across grounding, new information application, and knowledge acquisition using a 3,443-instance benchmark spanning science, finance, long documents, spatial reasoning, and web VQA; finds current multimodal models perform poorly (best score 0.2847) and analyzes failure modes.

Hugging Face

Provides 2,000 hours of synchronized, high‑fidelity robot‑free bimanual manipulation demonstrations with multi‑view video, calibrated end‑effector trajectories, gripper states, and language annotations. Curated from a 20,000+ hour corpus; features 6 camera views, ~3 mm pose accuracy, <40 µs cross‑sensor sync, and LeRobot v3‑style Parquet+MP4 export under CC BY 4.0.

Hugging Face

A collection of biology-focused 'mystery' tasks for benchmarking model performance on biomedical reasoning, evidence synthesis, and problem solving; curated by Anthropic and hosted on Hugging Face, designed for granular evaluation of scientific decision-making.