Human judgments of generated video typically focus on how things look or whether an explicit instruction is fulfilled. WorldExam pushes the evaluation frontier by treating controllable video generators as world models that must also exhibit inherent reactivity: given a scene state, infer plausible consequences the world should produce even when not explicitly commanded. This reframing reveals gaps that conventional visual-quality or instruction-following metrics miss.
Key Findings
- Hierarchical design: four diagnostic levels (Visual Quality, Control Adherence, Spatial Consistency, World Reactivity) and eight targeted tasks distill different failure modes so researchers can diagnose why a model fails to behave like a coherent world.
- Scale and scope: 1,474 curated cases across camera-, action-, and language-driven paradigms enable unified comparison across interface types rather than per-task isolated tests.
- Reactivity-focused metrics: World Reactivity tasks assess scene-conditioned reactions and goal-directed behavior beyond explicit inputs, exposing whether a model infers and generates plausible downstream consequences.
- Empirical split: evaluation of 20 representative models shows complementary strengths and weaknesses — camera-driven models handle camera control well but lack dynamic interaction; action-driven models control subjects precisely yet often leave the environment unresponsive; language-driven models handle interaction better but struggle with complex, fine-grained controls. No model combines broad task coverage with consistent reactivity.
Who it's for and tradeoffs
Great fit if you need a diagnostic benchmark to compare controllable video generators as world models, especially when your goal is to measure emergent, scene-conditioned behaviors rather than only visual fidelity or direct instruction following. It helps teams prioritize research on model-world coupling, interaction fidelity, and long-horizon consistency.
Look elsewhere if you only care about raw perceptual quality metrics or single-action instruction compliance — WorldExam emphasizes behavioral plausibility and cross-task diagnostics, which requires more annotation effort and is less focused on pixel-perfect perceptual scores. Use it alongside traditional quality benchmarks when you want to probe whether a model truly behaves like a reactive simulated world.