Most large pretraining crawls contain little explicit chain-of-thought or structured exercises, which slows small models' acquisition of reasoning skills. SYNTH tackles this by turning a curated encyclopedia core into a massively amplified collection of synthetic exercises paired with intermediary reasoning traces, so small models can learn reasoning patterns with far fewer training tokens.
What Sets It Apart
- Seeded amplification: built from 58,698 Wikipedia pages (Wikipedia:Vital Articles) plus targeted Wikibooks and internal samples, then amplified (≥100× on average) into ~79.65M synthetic samples. This focuses memorization and reduces reliance on unfocused web noise.
- Reasoning-by-design: every generated answer is accompanied by intermediary reasoning drafts in a reproducible syntax, making the dataset suitable for teaching chain-of-thought and RAG-style training workflows.
- Data-efficiency demonstrated: Pleias reports training small models (e.g., Monad 56M, Baguettotron family) to competitive benchmark performance using ≈100–200B tokens from SYNTH, markedly less data than typical web-scale mixtures.
- Practical structure and tooling: records include language, exercise type, generation constraints, query seed URL/text, synthetic reasoning and answer; data provided as parquet files for easy integration with common tooling.
Who it's for — and tradeoffs
Great fit if you want a reproducible, open dataset to pretrain or mid-train small-to-midsize reasoning models, to study how synthetic reasoning traces affect learning, or to run explainability/memorization experiments with controlled seeds. Look elsewhere if you need code-generation corpora (SYNTH intentionally excludes code), full global multilingual coverage beyond eight European languages, or training targets of many-billion-parameter models (SYNTH difficulty is calibrated for small models).
Where it fits
SYNTH is best viewed as a complementary, engineered alternative to generic web crawls when the goal is rapid, reproducible iteration on small reasoning models, or when you want explicit synthetic traces and provenance back to curated encyclopedic seeds.