AIAny
Icon for item

Fable-5 Premium Dataset

A cleaned supervised fine-tuning dataset of 6,365 Claude Fable-5 agent traces in OpenAI Chat and Hugging Face agent-traces formats, prepared for SFT, tool-use training, and distillation workflows; MIT-licensed and distributed as Parquet.

Introduction

Agent traces are uniquely valuable when you need training data that captures tool calls, multi-turn reasoning and structured assistant behavior. Fable-5 Premium prioritizes trace quality over quantity: every session is deduplicated, structurally validated, PII-scrubbed and scored for multi-dimensional quality so models learn from reliable SFT examples rather than noisy logs.

What Sets It Apart
  • Curated agent-trace focus — 6,365 premium Claude Fable-5 traces (after filtering/dedup) with explicit messages arrays and validated tool calls, so fine-tuned models get realistic assistant+tool interaction patterns rather than synthetic or truncated chat snippets.
  • SFT-ready formats — provided in OpenAI Chat (messages with user/assistant/tool roles) and native Hugging Face agent-traces formats (Parquet), so integration with Axolotl, Unsloth, OpenAI fine-tuning API or HF Data Studio is straightforward.
  • Rigorous quality pipeline — SHA-256 deduplication, schema validation, tool-response matching, PII scrubbing and multi-metric quality scores (average ~0.873) reduce noisy signal and lower the risk of models learning shortcuts or leaking sensitive artifacts.
  • Train/validation/test splits and metadata — pre-split data (train: 5,728 / val: 318 / test: 319), provenance fields, and quality annotations enable reproducible experiments and targeted selection of high-quality examples for distillation or targeted SFT.
Who It's For and Trade-offs

Great fit if you are training or fine-tuning assistant-style models that must call external tools, follow multi-turn flows, or benefit from high-integrity human-style traces; it’s also useful for distillation and tool-use evaluation. Look elsewhere if you need extremely large-scale raw web crawls for pretraining, multimodal data, or domain-specific labeled corpora (e.g., medical annotations) — this collection emphasizes quality of agent interactions over raw volume.

Information

  • Websitehuggingface.co
  • OrganizationsHugging Face, Crownelius collection
  • AuthorsSai Dutta Abhishek Dash
  • Published date2026/07/30

Categories

More Items

Hugging Face

Provides a machine-readable catalog of 117 AI/AX safety and deployment-readiness diagnostic criteria for assessing model intrinsic and serving/infrastructure risks. Includes MODEL-SCAN and AX-SCAN axes, bilingual source fields, per-item evidence guidance, severity/assurance metadata, and a CC BY-NC 4.0 release-candidate.

Hugging Face

Evaluates schema-guided structured extraction from documents: given a document and a JSON schema, systems must return a schema-valid JSON with page-and-box grounding. Covers 370 documents (4,869 pages) across 8 business domains and 67 document types; scores value accuracy, word/page grounding, and long-list completeness.

Hugging Face

Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.