Learns joint predictive visual dynamics and action generation for generalist robot manipulation, translating future-relevant visual representations into actions. Integrates a Mixture-of-Transformers coupling a video expert and action expert, a frozen vision-language model for semantics, 4D distillation, and Causal Imprint; pretrained on a 20K+ hour heterogeneous corpus.
Introduces an open-source multi-agent harness that automatically constructs, composes, and evolves modular model–harness units for long‑horizon, cross‑domain workflows. Key ingredients include a Host Agent for task decomposition and orchestration, EverOS for durable memory, and a Skill Forge of reusable procedures to improve task coverage via composition.