Provides 50 ARC‑AGI‑3 gameplay trajectories (GPT‑5.6 Sol and Claude Opus/Fable) plus a dependency‑free scorer and event logs; includes sanitized session data, snapshots, and utilities to recompute RHAE scores for reproducible agent evaluation and cross-model comparison.
Provides verified, model-attested end-to-end agent coding and debugging trajectories (JSONL). Each whole-session trace was produced by moonshotai/kimi-k3 on the pi/openrouter runtime, passed acceptance tests and independent model screening — useful for SFT, distillation, and analyzing tool-use behavior.