AIAny
Icon for item

HappyWorld-Bench

Evaluates whether generative world models maintain consistent, controllable, and physically plausible simulated environments under exploration, interaction, and intervention. Introduces a six-level W1–W6 capability taxonomy across three tracks (video, spatial, embodied) with human A/B Arena and automated metrics to measure behavioral correctness.

Introduction

Why this matters Most generative evaluation focuses on visual fidelity, but agents that plan or act inside generated worlds need those worlds to preserve state, respond predictably to actions, and support valid physical operations. The core insight of this work is that reliability under interaction is distinct from visual quality—HappyWorld-Bench stresses models with interaction-driven tests and human comparisons to reveal practical failure modes.

Key Findings
  • Unified taxonomy and multi-track protocol: defines six hierarchical capabilities (W1–W6) and operationalizes them across three evaluation tracks (video, spatial, embodied), so evaluations target perception, manipulability, and action-conditioned dynamics rather than only frame quality. This makes comparisons more directly relevant for planning and control use cases.
  • Human + automated evaluation pipeline: combines large-scale A/B human judgments (HappyWorld-Arena) with automated behavioral metrics, producing Elo ratings and fine-grained correctness scores — so you can see both subjective preference and objective state/causal failures.
  • Empirical gaps despite high fidelity: evaluated systems often score well visually but fail to preserve state across long rollouts, execute edits or placements reliably, and respond precisely to altered action conditions. This implies current generative models are limited as actionable internal simulators for downstream agents.
Who it's for and tradeoffs

Great fit if you need to benchmark or select generative world models for robotics, embodied AI, simulation-driven planning, or interactive content generation and care about state consistency and action-grounded correctness. The suite is useful for researchers comparing modeling choices (video rollouts, scene export, egocentric prediction) and for teams wanting human-calibrated preferences alongside automated metrics.

Look elsewhere if your main concern is pure single-frame photorealism or lightweight image/video synthesis benchmarking — HappyWorld-Bench prioritizes interaction fidelity and behavioral correctness, which requires more complex test cases and human evaluation overhead.

Where it fits

Positions itself between visual-quality benchmarks and task-driven simulators: not a drop-in physics engine, but a rigorous, behavioral evaluation layer for generative world models that must support exploration, interaction, and multi-step interventions.

Information

  • Websitearxiv.org
  • AuthorsZhiqi Bai, Junai Cai, Yixin Chen, Jingrun Du, Tao Feng, Wei Gong, Siyuan Huang, Xiao Lin, Jiaheng Liu, Jun Luo
  • Published date2026/09/21

More Items

Introduces a rubric-based approach for video reward modeling that generates explicit, query-adaptive evaluation criteria before scoring to reduce scalar drift; proposes RGPO, a two-stage training (seed warm-up + joint optimization) that achieves strong pointwise and pairwise evaluation with high data efficiency.

Teaches vision-language models to predict and integrate physical-world state transitions from interaction trajectories (observation → action → next observation). Introduces a three-level curriculum and the LSI-108K dataset, and applies supervised fine-tuning plus on-policy distillation to improve local transition modeling and long-horizon spatial integration.

An all-in-one multilingual scene text recognition approach that pairs a shared visual encoder with a script-aware Mixture-of-Experts (ScriptMoE) decoder to route each image to top-2 script experts. Introduces TextMuSS-10M, a synthetic dataset covering 10 scripts and 229 languages, and reports state-of-the-art accuracy and large end-to-end OCR F1 gains while remaining parameter-efficient.