Agent traces are uniquely valuable when you need training data that captures tool calls, multi-turn reasoning and structured assistant behavior. Fable-5 Premium prioritizes trace quality over quantity: every session is deduplicated, structurally validated, PII-scrubbed and scored for multi-dimensional quality so models learn from reliable SFT examples rather than noisy logs.
What Sets It Apart
- Curated agent-trace focus — 6,365 premium Claude Fable-5 traces (after filtering/dedup) with explicit
messagesarrays and validated tool calls, so fine-tuned models get realistic assistant+tool interaction patterns rather than synthetic or truncated chat snippets. - SFT-ready formats — provided in OpenAI Chat (
messageswith user/assistant/tool roles) and native Hugging Face agent-traces formats (Parquet), so integration with Axolotl, Unsloth, OpenAI fine-tuning API or HF Data Studio is straightforward. - Rigorous quality pipeline — SHA-256 deduplication, schema validation, tool-response matching, PII scrubbing and multi-metric quality scores (average ~0.873) reduce noisy signal and lower the risk of models learning shortcuts or leaking sensitive artifacts.
- Train/validation/test splits and metadata — pre-split data (train: 5,728 / val: 318 / test: 319), provenance fields, and quality annotations enable reproducible experiments and targeted selection of high-quality examples for distillation or targeted SFT.
Who It's For and Trade-offs
Great fit if you are training or fine-tuning assistant-style models that must call external tools, follow multi-turn flows, or benefit from high-integrity human-style traces; it’s also useful for distillation and tool-use evaluation. Look elsewhere if you need extremely large-scale raw web crawls for pretraining, multimodal data, or domain-specific labeled corpora (e.g., medical annotations) — this collection emphasizes quality of agent interactions over raw volume.