AIAny
Icon for item

Fable-5 Premium Dataset

A cleaned supervised fine-tuning dataset of 6,365 Claude Fable-5 agent traces in OpenAI Chat and Hugging Face agent-traces formats, prepared for SFT, tool-use training, and distillation workflows; MIT-licensed and distributed as Parquet.

Introduction

Agent traces are uniquely valuable when you need training data that captures tool calls, multi-turn reasoning and structured assistant behavior. Fable-5 Premium prioritizes trace quality over quantity: every session is deduplicated, structurally validated, PII-scrubbed and scored for multi-dimensional quality so models learn from reliable SFT examples rather than noisy logs.

What Sets It Apart
  • Curated agent-trace focus — 6,365 premium Claude Fable-5 traces (after filtering/dedup) with explicit messages arrays and validated tool calls, so fine-tuned models get realistic assistant+tool interaction patterns rather than synthetic or truncated chat snippets.
  • SFT-ready formats — provided in OpenAI Chat (messages with user/assistant/tool roles) and native Hugging Face agent-traces formats (Parquet), so integration with Axolotl, Unsloth, OpenAI fine-tuning API or HF Data Studio is straightforward.
  • Rigorous quality pipeline — SHA-256 deduplication, schema validation, tool-response matching, PII scrubbing and multi-metric quality scores (average ~0.873) reduce noisy signal and lower the risk of models learning shortcuts or leaking sensitive artifacts.
  • Train/validation/test splits and metadata — pre-split data (train: 5,728 / val: 318 / test: 319), provenance fields, and quality annotations enable reproducible experiments and targeted selection of high-quality examples for distillation or targeted SFT.
Who It's For and Trade-offs

Great fit if you are training or fine-tuning assistant-style models that must call external tools, follow multi-turn flows, or benefit from high-integrity human-style traces; it’s also useful for distillation and tool-use evaluation. Look elsewhere if you need extremely large-scale raw web crawls for pretraining, multimodal data, or domain-specific labeled corpora (e.g., medical annotations) — this collection emphasizes quality of agent interactions over raw volume.

Information

  • Websitehuggingface.co
  • OrganizationsHugging Face, Crownelius collection
  • AuthorsSai Dutta Abhishek Dash
  • Published date2026/07/30

Categories

More Items

Hugging Face

Provides 7,366 recorded agent trajectories from H Company’s Holo4 benchmark runs, with step-level reasoning, actions, tool results, token usage and screenshots for replay and analysis. Bundled as JSON and image files for per-trajectory inspection and automated replay; released under Apache 2.0.

Hugging Face

Provides 150,000 source‑grounded decision examples for training models that pick options, judge yes/no propositions, or assign ordered scores. Each row pairs a 'state' with JEV-style typed questions (CHOICE/NOUL/SCORE); multiple configs and train/test splits are included, license mixed/unknown.

Hugging Face

Provides 23,625 semi-structured smart-contract audit findings (title, description, PoC, recommendation, normalized severity) for defensive-security research; requires cleaning, deduplication, and PoC filtering before model training.