AIAny
Icon for item

bigfacing/GOKU-2M

Provides ~2 million instruction-aligned video-edit pairs for training and evaluating instruction-based video editing and generation models. Covers multi-task and structural edits (e.g., camera/subject movement), produced via a synthesis pipeline with progressive filtering; licensed CC BY-NC-4.0.

Introduction

Most public video-editing datasets focus on simple appearance swaps or single-task edits, which limits progress on instruction-driven, multi-step creative edits. This dataset addresses that gap by offering roughly two million instruction-aligned video editing pairs that explicitly include multi-task and structural manipulations, produced with automated synthesis and multi-stage quality filtering to improve alignment and temporal coherence.

What Sets It Apart
  • Large, instruction-aligned scale: ~2M edit pairs designed to support training large video-editing and text-to-video models and enable robust generalization across diverse edits.
  • Beyond appearance edits: includes basic appearance changes plus multi-task edits and structural transformations such as camera and subject movement, enabling researchers to tackle spatial and temporal control.
  • Scalable synthesis + progressive filtering: complex edits are decomposed into controllable subproblems, then filtered across instruction alignment, frame-to-frame stability, and perceptual realism to reduce noise from automated pipelines.
  • Benchmark & model ecosystem: paired with a human-verified Goku-Bench (1,000 test cases, 7 specialized metrics) and used to develop Goku-Edit (MLLM text encoder + dual-branch mask design), showing measurable gains in instruction following.
Who It's For and Trade-offs

Great fit if you need large-scale, instruction-aligned training data for research or prototyping of text-/instruction-to-video editing and video-to-video editing models, especially when structural edits and multi-task workflows are important. It is also useful for building or evaluating benchmarks for instruction-following in video editing.

Look elsewhere if you require purely real-capture, fully human-annotated edits without any synthetic data or if you need permissive commercial licensing—the dataset is released under CC BY-NC-4.0, which restricts commercial use. Automated synthesis and cascade processing can still introduce subtle artifacts despite progressive filtering, so downstream validation is recommended for high-stakes production use.

Information

  • Websitehuggingface.co
  • OrganizationsUniversity of Science and Technology of China, Tencent Hunyuan
  • AuthorsSen Liang, Cong Wang, Zhentao Yu, Fengbin Guan, Zhengguang Zhou, Teng Hu, Youliang Zhang, Yuan Zhou, Xin Li, Qinglin Lu …
  • Published date2026/06/22

Categories

More Items

Hugging Face

Provides a sanitized, labeled SOC capture and merged provenance graph for intrusion-detection research, including 2,011,674 live signals, 51,371 incident graphs, MITRE ATT&CK mappings, and deterministic attack reports.

Hugging Face

Aggregated, screened corpus of 55,050 normalized Indian public-information text bodies and 65,209 source records for retrieval and question-answering. Exports include deduplicated CSV/Parquet with provenance, topic labels, extraction quality flags and a private SQLite backup.

Hugging Face

Provides 3,451 hours (2,051,810 clips) of AI‑generated 48 kHz Turkish speech with transcripts, spoken forms and per‑clip voice descriptions for TTS and ASR development. Includes 2,752 designed voices and is licensed CC BY 4.0 / CC BY‑SA 4.0 (attribution to PatientDesk AI required).