AIAny
Icon for item

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

Introduces SemComp-Bench: a benchmark and VLM-based evaluation protocol for measuring outcome achievement and task-relevant semantic grounding in instruction-driven video generation. Ships with SemComp-Data, curated image–instruction–outcome triplets and OA/GR scoring.

Introduction

Most recent video generators prioritize visual fidelity and temporal coherence, but they rarely get evaluated on whether a generated clip actually achieves a specified outcome while staying semantically grounded in a reference image. This paper reframes the problem as Semantic Task Completion Video Generation and provides both data and a repeatable, interpretable evaluation protocol to measure that capability.

Key Findings
  • A focused evaluation target: the benchmark separates outcome achievement from low-level appearance fidelity and emphasizes whether the generated outcome matches the instructed goal and preserves task-relevant semantics from the reference. So what? Models can look realistic yet fail the task; this benchmark exposes that gap.
  • SemComp-Data design: constructs image–instruction–outcome triplets by mining full-context real videos and uses a four-stage curation pipeline (candidate filtering, state mining, video extension, instruction structuring). So what? Tasks are authentically achievable for each instance and preserve fine-grained task-relevant alignment.
  • VLM-based, evidence-grounded scoring: SemComp-Bench poses structured binary questions to a vision–language model and reports OA (Outcome Achievement) and GR (Generation Reliability) scores, with criterion-level pass rates for interpretable failure diagnosis. So what? This makes automated evaluation actionable and diagnostic rather than a single opaque metric.
  • Empirical gap: evaluations on representative video generators show substantial failures in completing instructed tasks while preserving reference-grounded semantics. So what? Progress in visual fidelity does not imply task competence; targeted research on outcome grounding is needed.
Who it's for and tradeoffs

Great fit if you evaluate or train video generation models where the objective is to realize user-specified outcomes (e.g., instruction-driven editing, simulation of object state changes). The benchmark is most useful for diagnostics and dataset-driven improvements rather than measuring pure perceptual quality. Look elsewhere if your primary concern is unconditional aesthetic quality, frame-level temporal realism without task semantics, or short synthetic clips that lack real-world task feasibility evidence.

Information

  • Websitearxiv.org
  • OrganizationsUniversity of Science and Technology of China, FrameX.AI, Sun Yat-sen University
  • AuthorsKeyu Tu, Zhuowei Chen, Mengqi Huang, Yuxin Wang, Jiahao Zhu, Zhendong Mao, Yongdong Zhang
  • Published date2026/08/18

More Items

Turns live video into reusable textual memory and timely responses by training a streaming video LLM to proactively generate time-grounded captions and event summaries. Key components: Proactive Hierarchical Caption Memory (PHCM) for multi-scale records and Proactive State Transition Learning (PSTL) to balance response timing; trained on the OneStreamer-1M streaming dataset.

Proposes a forward-process RL method that dynamically localizes gradient updates and adapts multi-reward coordination for joint audio–video diffusion models. Uses bidirectional cross-attention responses for token/layer routing and preference-preserving reweighting to balance competing objectives, improving modality quality, alignment, and synchronization.

Uses a multimodal model's own critiques as privileged context and applies on-policy self-distillation over diffusion sampling trajectories to internalize corrective guidance, improving text-to-image generation without an external teacher; shows measurable gains on GenEval and GenEval2.