Multimodal pretraining needs large, interleaved sequences of text and images — MINT-1T addresses that gap by scaling open-source multimodal data roughly 10× over previous public datasets. The HTML subset provides cleaned, filtered, and deduplicated HTML documents (data_v1_1), making it directly usable as a large training shard for models that learn from interleaved image/text sequences.
What Sets It Apart
- Scale and diversity: the full MINT-1T collection totals ≈1.0 trillion text tokens and ≈3.4 billion images, sourced from HTML, PDFs, and ArXiv papers; the HTML subset (data_v1_1) contains ~742B HTML tokens after boilerplate removal and additional safety filtering. This is an order-of-magnitude increase compared to prior open datasets like OBELICS (~115B tokens).
- Safety- and quality-focused curation: multi-stage image filtering (NSFW classifier + Datacomp safety classifier), image size/aspect-ratio thresholds, language identification, PII masking, paragraph- and document-level deduplication, and PDF/ArXiv-specific parsing to preserve reading order and figure/text interleaving.
- Interleaved format: documents preserve free-form sequences of text and images rather than separate image/text corpora, making the data directly applicable to training models that ingest interleaved multimodal streams (e.g., Idefics2, XGen-MM, Chameleon).
Who It’s For + Tradeoffs
- Great fit if you need very large-scale open multimodal pretraining data for research or prototyping multimodal LMMs and value interleaved image/text context at web scale.
- Look elsewhere if you require guaranteed absence of any personal data or copyrighted media for commercial deployment: despite masking and filtering, the corpus is drawn from public web crawls and may still contain sensitive or copyrighted content. Users should perform additional filtering tailored to their legal and ethical constraints.
Practical notes
- Licensing: released under CC-BY-4.0 (research-focused; verify commercial/legal compliance before commercial use).
- Recommended workflow: treat MINT-1T HTML shards as pretraining material after applying project-specific additional filters, sample balancing, and checks for image availability (link rot can affect reproducibility).