Why this matters Most generative evaluation focuses on visual fidelity, but agents that plan or act inside generated worlds need those worlds to preserve state, respond predictably to actions, and support valid physical operations. The core insight of this work is that reliability under interaction is distinct from visual quality—HappyWorld-Bench stresses models with interaction-driven tests and human comparisons to reveal practical failure modes.
Key Findings
- Unified taxonomy and multi-track protocol: defines six hierarchical capabilities (W1–W6) and operationalizes them across three evaluation tracks (video, spatial, embodied), so evaluations target perception, manipulability, and action-conditioned dynamics rather than only frame quality. This makes comparisons more directly relevant for planning and control use cases.
- Human + automated evaluation pipeline: combines large-scale A/B human judgments (HappyWorld-Arena) with automated behavioral metrics, producing Elo ratings and fine-grained correctness scores — so you can see both subjective preference and objective state/causal failures.
- Empirical gaps despite high fidelity: evaluated systems often score well visually but fail to preserve state across long rollouts, execute edits or placements reliably, and respond precisely to altered action conditions. This implies current generative models are limited as actionable internal simulators for downstream agents.
Who it's for and tradeoffs
Great fit if you need to benchmark or select generative world models for robotics, embodied AI, simulation-driven planning, or interactive content generation and care about state consistency and action-grounded correctness. The suite is useful for researchers comparing modeling choices (video rollouts, scene export, egocentric prediction) and for teams wanting human-calibrated preferences alongside automated metrics.
Look elsewhere if your main concern is pure single-frame photorealism or lightweight image/video synthesis benchmarking — HappyWorld-Bench prioritizes interaction fidelity and behavioral correctness, which requires more complex test cases and human evaluation overhead.
Where it fits
Positions itself between visual-quality benchmarks and task-driven simulators: not a drop-in physics engine, but a rigorous, behavioral evaluation layer for generative world models that must support exploration, interaction, and multi-step interventions.