Robotic manipulation often breaks down because predictive visual models are not directly usable for control: generating future video rollouts is expensive and predictive features may not align with what actions need. InternW0-Δ flips that equation by training predictive visual dynamics specifically to produce future-relevant representations that an action policy can consume without sampling future frames at inference. This makes long-horizon, heterogeneous pretraining actually actionable for downstream robot controllers.
Key Findings
- Joint video–action pretraining with a directed Mixture-of-Transformers lets a high-capacity video expert provide episodic and recent-context predictions while a lightweight action expert runs at control time, so what: the system reuses predictive context without expensive per-step rollout.
- Causal Imprint learns future-relevant scene changes via training-only supervision and exposes those representations to the action expert, so what: the policy benefits from predictive cues without accessing future observations at inference.
- 4D-aware distillation injects geometric and motion priors from a pretrained 4D foundation model into the video expert during training, so what: improves spatial and motion consistency for manipulation tasks that require geometry-aware reasoning.
- Large heterogeneous corpus (20K+ hours) spanning robot demos, egocentric human data, and Ego2Robot-style conversions improves generalization across embodiments and benchmarks; reported gains include top performance on multiple simulation suites and real-robot adaptations.
Who it's for and trade-offs
Great fit if you need a single pretrained world-action model to adapt to multiple robot platforms and tasks, especially when you can afford large-scale pretraining and post-training for target embodiments. Look elsewhere if you require ultra-low-latency on-device inference without model adaptation, or if you cannot provide compute/resources for post-training or finetuning on target controllers. The architecture emphasizes representational reuse and cross-modal priors, which increases pretraining complexity and model size compared to simple imitation baselines.
How it works (brief)
The system couples: a pretrained video expert that consumes sparse anchor/recent/current frames for predictive context; an action expert that consumes scene semantics from a frozen vision-language model and Causal Imprint features; and a training-only 4D distillation branch that supplies geometric/motion priors. Directed information flow prevents future observations from being used at inference; future data is used only as supervision during training. Post-training adapts the shared checkpoint to specific control interfaces and embodiments.