AIAny
Icon for item

ConceptEdit-12M

Provides 12 million verified source/edited image pairs with per-sample edit instructions and VQA-style quality checks for large-scale training and evaluation of instruction-based image editing models. Features a 1,000+ fine-grained edit taxonomy and multi-concept dense-supervision bundles; data is distributed as TAR shards for scalable extraction.

Introduction

Most instruction-based image editing efforts are constrained by limited, sparse supervision and coarse edit taxonomies. ConceptEdit-12M addresses both by supplying a massive, taxonomy-grounded dataset plus dense multi-concept supervision to increase signal per training example and improve edit granularity.

What Sets It Apart
  • Scale and verification: 12 million verified (source, edited, metadata) triplets with VQA-style checks for edit correctness — so what: enables training at scales comparable to large T2I corpora while filtering low-quality synthetic edits.
  • Fine-grained taxonomy: a hierarchical edit concept library spanning 1,000+ edit categories — so what: supports targeted instruction tuning and granular evaluation across diverse editing intentions.
  • Dense supervision variants: bundles multiple non-interfering edits into single paired examples — so what: supplies richer per-sample learning signals, improving training efficiency and multi-edit competence.
  • Practical packaging: released as four main split folders of TAR shards with relative paths and per-sample JSONs — so what: facilitates batch extraction and integration into distributed training pipelines.
Who It's For and Tradeoffs

Great fit if you train or benchmark instruction-conditioned image editing models (diffusion-based editors, instruction-tuned i2i models) and need large-scale, taxonomy-aware synthetic supervision. Look elsewhere if you require exclusively human-labeled real-world edits or proprietary image sources: ConceptEdit’s edits are synthesized and its source images derive from Fine-T2I; modelers should validate applicability to their real-world distribution. The dataset is released under Apache-2.0 and packaged for scalable extraction, but working with 12M examples implies significant storage and compute costs.

Information

  • Websitehuggingface.co
  • OrganizationsinclusionAI
  • AuthorsLong Cui, Xiaoqian Liu, Qi Qin, Yi Xin, Tao Lin, Jianguo Li, Linfeng Zhang
  • Published date2026/08/17

Categories

More Items

Hugging Face

Curated English Wikipedia text prepared for language-model training and evaluation, provided in WikiText-2 and WikiText-103 variants. Preserves original case, punctuation and numbers; offers raw and tokenized splits for long-range language modeling under a CC BY‑SA license.

Hugging Face

Provides 16 weeks of anonymized production agent-session traces (12,002 sessions, ~1.19M LLM requests, ~1.21M tool calls, 209B input tokens) released as block-level prefix IDs plus flattened Parquet tables for KV-cache, scheduling and serving-system research.

Hugging Face

Open synthetic corpus for training small reasoning-focused language models — ~79.65M generated samples (≈75B tokens with Pleias tokenizer) amplified from ~58.7k Wikipedia/Wikibooks seeds; includes explicit synthetic reasoning traces, multilingual coverage, and parquet splits.