Uses large-scale, mixed-task reinforcement learning to drive self-improvement of multimodal foundation models. Key features include fully-asynchronous large-batch RL (1,568 prompts × 16 rollouts, up to 3.7B tokens/update, 1M context), groupwise agentic grading, and sparse MoE architectures.