AIAny
Icon for item

SYNTH - generalist open data and environment

Open synthetic corpus for training small reasoning-focused language models — ~79.65M generated samples (≈75B tokens with Pleias tokenizer) amplified from ~58.7k Wikipedia/Wikibooks seeds; includes explicit synthetic reasoning traces, multilingual coverage, and parquet splits.

Introduction

Most large pretraining crawls contain little explicit chain-of-thought or structured exercises, which slows small models' acquisition of reasoning skills. SYNTH tackles this by turning a curated encyclopedia core into a massively amplified collection of synthetic exercises paired with intermediary reasoning traces, so small models can learn reasoning patterns with far fewer training tokens.

What Sets It Apart
  • Seeded amplification: built from 58,698 Wikipedia pages (Wikipedia:Vital Articles) plus targeted Wikibooks and internal samples, then amplified (≥100× on average) into ~79.65M synthetic samples. This focuses memorization and reduces reliance on unfocused web noise.
  • Reasoning-by-design: every generated answer is accompanied by intermediary reasoning drafts in a reproducible syntax, making the dataset suitable for teaching chain-of-thought and RAG-style training workflows.
  • Data-efficiency demonstrated: Pleias reports training small models (e.g., Monad 56M, Baguettotron family) to competitive benchmark performance using ≈100–200B tokens from SYNTH, markedly less data than typical web-scale mixtures.
  • Practical structure and tooling: records include language, exercise type, generation constraints, query seed URL/text, synthetic reasoning and answer; data provided as parquet files for easy integration with common tooling.
Who it's for — and tradeoffs

Great fit if you want a reproducible, open dataset to pretrain or mid-train small-to-midsize reasoning models, to study how synthetic reasoning traces affect learning, or to run explainability/memorization experiments with controlled seeds. Look elsewhere if you need code-generation corpora (SYNTH intentionally excludes code), full global multilingual coverage beyond eight European languages, or training targets of many-billion-parameter models (SYNTH difficulty is calibrated for small models).

Where it fits

SYNTH is best viewed as a complementary, engineered alternative to generic web crawls when the goal is rapid, reproducible iteration on small reasoning models, or when you want explicit synthetic traces and provenance back to curated encyclopedic seeds.

Information

  • Websitehuggingface.co
  • OrganizationsPleIAs, AI Alliance, Wikimedia Foundation / Wikimedia Enterprise, Wikipedia community (Wikipedia:Vital Articles)
  • Published date2025/11/10

Categories

More Items

Hugging Face

De-identified longitudinal multimodal CT dataset for multicancer screening that pairs ~24k CT volumes with radiology reports and voxel-wise tumor annotations across 13 cancer types. Designed for longitudinal disease modeling, detection/segmentation and vision–language research; CC BY‑NC‑ND 4.0 for non-commercial use.

Hugging Face

Provides 98,877 historical newspaper page images (1700s–1940s) paired with ALTO OCR, per-word confidences and line/word bounding boxes in image pixels — ready for line-level OCR training, OCR quality estimation, re‑OCR comparisons and layout analysis. OCR is library-produced (silver); image resolutions and OCR quality vary.

Hugging Face

Provides 1.21M densely annotated desktop screenshots and 159.7M element instances for training and evaluating GUI grounding and screen-parsing models. Includes per-element accessibility-derived annotations, 917K recorded click transitions, multi-application scenes across seven appearance presets and resolutions; distributed as WebDataset shards with Parquet indexes.