AIAny
Icon for item

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

Uses large-scale, mixed-task reinforcement learning to drive self-improvement of multimodal foundation models. Key features include fully-asynchronous large-batch RL (1,568 prompts × 16 rollouts, up to 3.7B tokens/update, 1M context), groupwise agentic grading, and sparse MoE architectures.

Introduction

Most progress in foundation models focuses on scaling parameters or data; this work argues that scaling RL compute, environments, and grader capacity together creates a practical path toward recursive self-improvement for multimodal models. The concrete payoff is not just higher benchmark scores but a closed-loop that favors shorter, more token-efficient solutions learned from the model’s own rollouts.

Key Findings
  • Large asynchronous batches and long contexts: the training pipeline runs fully-asynchronous Group Relative Policy Optimization on huge batches (1,568 prompts × 16 rollouts per update), consuming billions of tokens per update (reported 2.7–3.7B) and supporting context lengths up to 1M tokens. This expands the effective exploration space for agentic behaviors and long-horizon tasks.
  • Groupwise agentic grading (GRS/GAR): instead of binary pass/fail rewards, the paper scales grader compute by comparing rollouts within groups to produce task-specific rubrics and rank passing trajectories. That yields richer, more accurate reward signals and steers policies toward shorter, token-efficient solutions.
  • Mixed-task, multi-harness RL: a single mixed RL run combines coding, general agents, visual, and cybersecurity tasks so capabilities can transfer across domains without separate per-domain RL runs.
  • Architecture and stability practices: uses sparse Mixture-of-Experts (MoE) backbones (e.g., a 1.02T total / 42B active Pro variant), freezes MoE routers during RL, and applies defenses (adversarial screening, verifiers) to mitigate reward-hacking during large-scale RL.
  • Post-RL distillation and verification: introduces Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2) to consolidate mixed-RL gains into deployable single-turn behaviors and reuses teacher prefixes to avoid regenerating long trajectories.
Who it's for and tradeoffs

Great fit if you are researching large-scale agentic RL, multimodal long-horizon behaviors, or infrastructure for high-throughput RL experiments and want concrete engineering patterns (asynchronous large-batch RL, groupwise grading, MoE stability measures). Look elsewhere if your resources are limited: the methods assume massive grader and compute budgets, complex environment/harness engineering, and nontrivial infrastructure to run 1M-token contexts and billions-of-tokens updates. The paper is valuable for system designers and RL researchers but not directly applicable as an off-the-shelf recipe for small labs.

Information

  • Websitearxiv.org
  • OrganizationsXiaomi LLM-Core Team, Xiaomi
  • AuthorsZongming Qiao, Ziyue Hua, Zirui Ou, Zihao Yue, Zihan Jiang, Zhuo Huang, Zhiyang Chen, Zhixian Zheng, Zhipeng Xu, Zhengrui Ma …
  • Published date2026/10/08

More Items

Proposes TRACE, an FP4 quantization framework for RL of MoE LLMs that uses rollout-side FP4 outcomes to guide training-side rounding and caches mantissa/scale from deeper layers to limit overhead. Preserves BF16-level RL performance while enabling up to 5.4× rollout speedup.

Adapts LLM agents online by self-distilling verified execution trajectories into persistent LoRA weights during deployment to improve success and efficiency on long‑horizon tasks. Uses a frozen stable copy as a privileged teacher to predict hindsight next‑token distributions and filters invalid-action turns so experience consolidates without external solutions or memory retrieval.

Calibrates multi-reward reinforcement learning by adaptively upweighting infrequently active rewards per rollout batch, so sparse objectives provide stronger signals when they matter. Proposes an inverse-square-root density correction and shows faster learning on tool-calling and math-reasoning tasks.