AIAny
Icon for item

Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction

Generates synchronized egocentric video streams for multiple interacting agents in a shared environment, enforcing cross-view action consistency, shared environment memory, and consistent propagation of interaction-induced state changes — aimed at embodied AI, VR, and multi-agent vision research.

Introduction

Multi-agent embodied settings force a world model to do more than predict a single first-person stream: an action by one agent must both look plausible from that agent’s view and be observed consistently by others, while shared objects and scene state evolve coherently. This paper argues that treating multiple ego streams as coupled observations of one evolving world — rather than independent generation targets — is the key to consistent multi-view embodied video prediction.

Key Findings
  • Joint generation beats independent per-agent models: denoising all ego streams as a single shared token sequence reduces contradictions across views, so interaction outcomes (moved objects, raised hands) appear consistently to every agent.
  • Shared action conditioning preserves cross-view action fidelity: conditioning each target stream on all agents’ target-view poses makes body and hand motions plausible both in the actor’s own stream and in other agents’ observations, so generated actions remain synchronized across viewpoints.
  • Shared environment memory grounds appearance and geometry: warping historical observations into each target view and using common anchor references maintains consistent scene appearance and propagates interaction-induced state updates across streams, improving environment and update consistency metrics.
  • Practical gains: the approach improves shared-world consistency, action-control alignment, identity preservation, and standard video quality metrics on both real and synthetic multi-agent benchmarks, with ablations showing each core component contributes meaningfully.
Who it's for + tradeoffs

Great fit if you research embodied AI, multi-agent vision, VR multi-user simulation, or multi-robot collaboration and need generative models that keep multiple first-person views consistent. It helps when fine-grained embodied actions (body, hand, head motion) must be synchronized and when interaction effects must propagate across views. Look elsewhere if your domain only requires single-view prediction, coarse high-level controls (e.g., camera trajectories or discrete commands), or very low-latency control loops where diffusion-based video generation is impractical. The method relies on joint denoising and environment memory, which increases model complexity and inference cost compared to per-view baselines.

Where it fits

ME-World sits between single-agent egocentric world models and multi-view generation methods that assume coarse controls: it targets scenarios where actions are both first-person controls and cross-view observable events, such as collaborative cooking, multi-user VR, or multi-robot manipulation.

Method overview

The architecture fine-tunes a video diffusion transformer to denoise all agents’ ego streams as one shared token sequence, supplies shared action conditioning (each stream gets its own viewing rays plus all agents’ body motions), and grounds generation with a shared environment memory built from all agents’ observation history. New shared-world consistency metrics (S_env, S_update, S_id on synthetic data) measure environment agreement, propagation of interaction-induced updates, and identity preservation across generated streams.

Information

  • Websitearxiv.org
  • OrganizationsAffiliation: Hyunsung Kim, Affiliation: [4pt] † Co-corresponding authors, Affiliation: [4pt] KAIST AI, Affiliation: [3pt] https://cvlab-kaist.github.io/ME-World
  • AuthorsDahyun Chung, Siyoon Jin, Hyunwook Choi, Honggyu An, Junyoung Seo, Hyunsung Kim, Seung Wook Kim, Seungryong Kim
  • Published date2026/10/08

More Items

Predicts identity-preserving dense pixel correspondences between image pairs that violate spatio-temporal priors (e.g., edits and reference-guided generation). Fuses generative (FLUX2) and semantic (DINOv3) foundation representations with heterogeneous supervision and teacher-guided iterative refinement to generalize beyond classical optical-flow assumptions.

Converts static 3D Gaussian Splatting scenes into endlessly looping 3D cinemagraphs by inferring plausible dynamics with a vision-language model, synthesizing a reference video, lifting it to multi-view videos, and fitting a Fourier-parameterized Periodic Deformation Field with a Grounded Drift Field—mask-free capture of deformation, object motion, and illumination changes.

Lets a pretrained multimodal LLM interpret navigation requests and orchestrate motion via tool calls for generalist robot navigation across unfamiliar scenes. Key features: an agent harness with Navigation Skills, a unified visual-point interface, task-progress tracking, and tool-based motion execution without navigation-specific fine-tuning.