Multi-agent embodied settings force a world model to do more than predict a single first-person stream: an action by one agent must both look plausible from that agent’s view and be observed consistently by others, while shared objects and scene state evolve coherently. This paper argues that treating multiple ego streams as coupled observations of one evolving world — rather than independent generation targets — is the key to consistent multi-view embodied video prediction.
Key Findings
- Joint generation beats independent per-agent models: denoising all ego streams as a single shared token sequence reduces contradictions across views, so interaction outcomes (moved objects, raised hands) appear consistently to every agent.
- Shared action conditioning preserves cross-view action fidelity: conditioning each target stream on all agents’ target-view poses makes body and hand motions plausible both in the actor’s own stream and in other agents’ observations, so generated actions remain synchronized across viewpoints.
- Shared environment memory grounds appearance and geometry: warping historical observations into each target view and using common anchor references maintains consistent scene appearance and propagates interaction-induced state updates across streams, improving environment and update consistency metrics.
- Practical gains: the approach improves shared-world consistency, action-control alignment, identity preservation, and standard video quality metrics on both real and synthetic multi-agent benchmarks, with ablations showing each core component contributes meaningfully.
Who it's for + tradeoffs
Great fit if you research embodied AI, multi-agent vision, VR multi-user simulation, or multi-robot collaboration and need generative models that keep multiple first-person views consistent. It helps when fine-grained embodied actions (body, hand, head motion) must be synchronized and when interaction effects must propagate across views. Look elsewhere if your domain only requires single-view prediction, coarse high-level controls (e.g., camera trajectories or discrete commands), or very low-latency control loops where diffusion-based video generation is impractical. The method relies on joint denoising and environment memory, which increases model complexity and inference cost compared to per-view baselines.
Where it fits
ME-World sits between single-agent egocentric world models and multi-view generation methods that assume coarse controls: it targets scenarios where actions are both first-person controls and cross-view observable events, such as collaborative cooking, multi-user VR, or multi-robot manipulation.
Method overview
The architecture fine-tunes a video diffusion transformer to denoise all agents’ ego streams as one shared token sequence, supplies shared action conditioning (each stream gets its own viewing rays plus all agents’ body motions), and grounds generation with a shared environment memory built from all agents’ observation history. New shared-world consistency metrics (S_env, S_update, S_id on synthetic data) measure environment agreement, propagation of interaction-induced updates, and identity preservation across generated streams.