Most progress in foundation models focuses on scaling parameters or data; this work argues that scaling RL compute, environments, and grader capacity together creates a practical path toward recursive self-improvement for multimodal models. The concrete payoff is not just higher benchmark scores but a closed-loop that favors shorter, more token-efficient solutions learned from the model’s own rollouts.
Key Findings
- Large asynchronous batches and long contexts: the training pipeline runs fully-asynchronous Group Relative Policy Optimization on huge batches (1,568 prompts × 16 rollouts per update), consuming billions of tokens per update (reported 2.7–3.7B) and supporting context lengths up to 1M tokens. This expands the effective exploration space for agentic behaviors and long-horizon tasks.
- Groupwise agentic grading (GRS/GAR): instead of binary pass/fail rewards, the paper scales grader compute by comparing rollouts within groups to produce task-specific rubrics and rank passing trajectories. That yields richer, more accurate reward signals and steers policies toward shorter, token-efficient solutions.
- Mixed-task, multi-harness RL: a single mixed RL run combines coding, general agents, visual, and cybersecurity tasks so capabilities can transfer across domains without separate per-domain RL runs.
- Architecture and stability practices: uses sparse Mixture-of-Experts (MoE) backbones (e.g., a 1.02T total / 42B active Pro variant), freezes MoE routers during RL, and applies defenses (adversarial screening, verifiers) to mitigate reward-hacking during large-scale RL.
- Post-RL distillation and verification: introduces Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2) to consolidate mixed-RL gains into deployable single-turn behaviors and reuses teacher prefixes to avoid regenerating long trajectories.
Who it's for and tradeoffs
Great fit if you are researching large-scale agentic RL, multimodal long-horizon behaviors, or infrastructure for high-throughput RL experiments and want concrete engineering patterns (asynchronous large-batch RL, groupwise grading, MoE stability measures). Look elsewhere if your resources are limited: the methods assume massive grader and compute budgets, complex environment/harness engineering, and nontrivial infrastructure to run 1M-token contexts and billions-of-tokens updates. The paper is valuable for system designers and RL researchers but not directly applicable as an off-the-shelf recipe for small labs.