AIAny
Icon for item

MINT-1T (HTML)

Provides a 1 trillion-token multimodal interleaved dataset (HTML subset updated as data_v1_1 with 742B HTML tokens) and 3.4B images drawn from HTML/PDF/ArXiv sources for multimodal pretraining; released under CC-BY-4.0 with safety and deduplication guidance.

Introduction

Multimodal pretraining needs large, interleaved sequences of text and images — MINT-1T addresses that gap by scaling open-source multimodal data roughly 10× over previous public datasets. The HTML subset provides cleaned, filtered, and deduplicated HTML documents (data_v1_1), making it directly usable as a large training shard for models that learn from interleaved image/text sequences.

What Sets It Apart
  • Scale and diversity: the full MINT-1T collection totals ≈1.0 trillion text tokens and ≈3.4 billion images, sourced from HTML, PDFs, and ArXiv papers; the HTML subset (data_v1_1) contains ~742B HTML tokens after boilerplate removal and additional safety filtering. This is an order-of-magnitude increase compared to prior open datasets like OBELICS (~115B tokens).
  • Safety- and quality-focused curation: multi-stage image filtering (NSFW classifier + Datacomp safety classifier), image size/aspect-ratio thresholds, language identification, PII masking, paragraph- and document-level deduplication, and PDF/ArXiv-specific parsing to preserve reading order and figure/text interleaving.
  • Interleaved format: documents preserve free-form sequences of text and images rather than separate image/text corpora, making the data directly applicable to training models that ingest interleaved multimodal streams (e.g., Idefics2, XGen-MM, Chameleon).
Who It’s For + Tradeoffs
  • Great fit if you need very large-scale open multimodal pretraining data for research or prototyping multimodal LMMs and value interleaved image/text context at web scale.
  • Look elsewhere if you require guaranteed absence of any personal data or copyrighted media for commercial deployment: despite masking and filtering, the corpus is drawn from public web crawls and may still contain sensitive or copyrighted content. Users should perform additional filtering tailored to their legal and ethical constraints.
Practical notes
  • Licensing: released under CC-BY-4.0 (research-focused; verify commercial/legal compliance before commercial use).
  • Recommended workflow: treat MINT-1T HTML shards as pretraining material after applying project-specific additional filters, sample balancing, and checks for image availability (link rot can affect reproducibility).

Information

  • Websitehuggingface.co
  • OrganizationsUniversity of Washington, Salesforce Research, Stanford University, University of Texas at Austin, University of California, Berkeley, mlfoundations
  • AuthorsAnas Awadalla, Le Xue, Oscar Lo, Manli Shu, Hannah Lee, Etash Kumar Guha, Matt Jordan, Sheng Shen, Mohamed Awadalla, Silvio Savarese …
  • Published date2024/07/21

Categories

More Items

Hugging Face

Provides 369 Harbor sandbox tasks ported from OpenAI's openai/math: each task is a Lean theorem with missing `sorry` proofs that an agent must complete, graded by a strict Comparator exact-match verifier. Includes task definitions, generator, and manifest for RL/code-agent evaluation.

Hugging Face

Curated English Wikipedia text prepared for language-model training and evaluation, provided in WikiText-2 and WikiText-103 variants. Preserves original case, punctuation and numbers; offers raw and tokenized splits for long-range language modeling under a CC BY‑SA license.

Hugging Face

Provides 16 weeks of anonymized production agent-session traces (12,002 sessions, ~1.19M LLM requests, ~1.21M tool calls, 209B input tokens) released as block-level prefix IDs plus flattened Parquet tables for KV-cache, scheduling and serving-system research.