Predicts future video frames conditioned on an observed frame, a language instruction, and a sequence of end-effector poses and gripper states for robot manipulation. Uses per-arm SE(3) geometric encoding (PRoPE-style), a lightweight depth branch, SAM3 masks with a frozen V-JEPA teacher, and distribution-matching distillation for efficient, consistent action-conditioned rollouts.
Provides a Mixture-of-Experts language model tuned for million-token contexts and agentic workflows, with DSpark speculative decoding, FP4/FP8 mixed-precision support, and vLLM/SGLang deployment recipes for low-latency production inference.