AIAny
Icon for item

InternW0-$Δ$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

Learns joint predictive visual dynamics and action generation for generalist robot manipulation, translating future-relevant visual representations into actions. Integrates a Mixture-of-Transformers coupling a video expert and action expert, a frozen vision-language model for semantics, 4D distillation, and Causal Imprint; pretrained on a 20K+ hour heterogeneous corpus.

Introduction

Robotic manipulation often breaks down because predictive visual models are not directly usable for control: generating future video rollouts is expensive and predictive features may not align with what actions need. InternW0-Δ flips that equation by training predictive visual dynamics specifically to produce future-relevant representations that an action policy can consume without sampling future frames at inference. This makes long-horizon, heterogeneous pretraining actually actionable for downstream robot controllers.

Key Findings
  • Joint video–action pretraining with a directed Mixture-of-Transformers lets a high-capacity video expert provide episodic and recent-context predictions while a lightweight action expert runs at control time, so what: the system reuses predictive context without expensive per-step rollout.
  • Causal Imprint learns future-relevant scene changes via training-only supervision and exposes those representations to the action expert, so what: the policy benefits from predictive cues without accessing future observations at inference.
  • 4D-aware distillation injects geometric and motion priors from a pretrained 4D foundation model into the video expert during training, so what: improves spatial and motion consistency for manipulation tasks that require geometry-aware reasoning.
  • Large heterogeneous corpus (20K+ hours) spanning robot demos, egocentric human data, and Ego2Robot-style conversions improves generalization across embodiments and benchmarks; reported gains include top performance on multiple simulation suites and real-robot adaptations.
Who it's for and trade-offs

Great fit if you need a single pretrained world-action model to adapt to multiple robot platforms and tasks, especially when you can afford large-scale pretraining and post-training for target embodiments. Look elsewhere if you require ultra-low-latency on-device inference without model adaptation, or if you cannot provide compute/resources for post-training or finetuning on target controllers. The architecture emphasizes representational reuse and cross-modal priors, which increases pretraining complexity and model size compared to simple imitation baselines.

How it works (brief)

The system couples: a pretrained video expert that consumes sparse anchor/recent/current frames for predictive context; an action expert that consumes scene semantics from a frozen vision-language model and Causal Imprint features; and a training-only 4D distillation branch that supplies geometric/motion priors. Directed information flow prevents future observations from being used at inference; future data is used only as supervision during training. Post-training adapts the shared checkpoint to specific control interfaces and embodiments.

Information

  • Websitearxiv.org
  • OrganizationsShanghai AI Laboratory: Affiliation: Project page:https://internrobotics.github.io/InternW0-Delta/
  • AuthorsXingyu Miao, Zizun Li, Baole Fang, Kaiwen Song, Tenghui Wang, Hanxue Zhang, Yating Wang, Xudong Li, Yuping He, Xueyuan Wei …
  • Published date2026/09/25

More Items

Uses diffusion-model forking moments as a proxy for perceptual distance to automatically generate pointwise reference-grounded labels, enabling annotation-free training of reference-based image quality assessment metrics.

Trains decoders and diffusion generators to be robust to random subsets of pretrained encoder layers by regularizing layer-fusion during training, narrowing the reconstruction–generation gap in representation autoencoders and improving ImageNet-256 PSNR and gFID.

Alternates a Planner (issues sub-queries) and a Synthesizer (integrates retrieved evidence into a persistent summary) to tackle long-horizon deep-search; introduces Role‑Decoupled Policy Optimization (RDPO) for role-specific RL credit assignment and shows strong results (IterSynth-8B reaches 50.7% on five benchmarks).