Why this matters
AI video generators can fabricate realistic depictions of wars, disasters and other crises that spread rapidly on social platforms. Detecting such videos is not just an artifact-matching problem: detectors must generalize across diverse generators, conditioning methods, and the social processes that amplify misinformation. This paper provides a focused empirical lens on those real-world challenges by building RA-Bench and evaluating detectors, generators, and dissemination together.
Key Findings
-
RA-Bench and evaluation design: introduces RA-Bench with 17,886 videos (1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips) produced by four open-source and five closed-source generators. The benchmark organizes analysis along three dimensions: detector generalization, generation-condition sensitivity, and social dissemination effects.
-
Detector generalization is weak: evaluates seven traditional detectors, ten zero-shot multimodal models under three review settings, and two MLLMs fine-tuned for generated-video detection. None of these families generalizes consistently across RA-Bench instances; performance varies substantially by generator and scenario.
-
Generation properties matter but patterns are stable by source: detectability depends on generation quality, conditioning information, and sampling seeds; however, source-level detection patterns remain relatively stable across seeds, implying generator-level biases.
-
Human judgments and dissemination interact with detectors: videos that frequently mislead humans are also harder for current detectors, and detection reliability degrades further after simulated social dissemination, highlighting a gap between lab benchmarks and field risk.
-
Practical implication: improving robustness requires evidence-based, semantics-aware verification (beyond low-level artifacts), broader generator coverage in training/benchmarks, and evaluation under realistic social pipelines.
Who it's for and tradeoffs
Great fit if you are a researcher or practitioner building or evaluating AI-generated video detectors, a dataset/benchmark developer aiming to capture real-world risk categories, or a policymaker assessing platform-level defenses. The study provides concrete failure modes (generator-dependent gaps, dissemination-induced degradation) and a sizable public benchmark to reproduce analyses.
Look elsewhere if you need a narrow artifact detector tuned to a specific generator or if your threat model excludes social dissemination: RA-Bench emphasizes cross-generator realism and social-context evaluation, so its findings may underrepresent scenarios limited to a single, well-known generator. The benchmark is comprehensive across many generators and categories but remains constrained to the included models, categories, and the authors' dissemination simulations.