AIAny
Icon for item

AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

Defines and evaluates AREX-2, an LLM agent that iteratively self-improves at test time via reflection and long-horizon execution. Trained on long-horizon improvement trajectories from ML engineering and algorithmic programming (built on Qwen3.8-27B), it scales with more rounds and achieves strong benchmark scores.

Introduction

Hard improvements in research and engineering rarely come from a single pass; they come from repeated try-measure-revise cycles. AREX-2 moves that loop inside the model: instead of asking the harness to orchestrate retries, it trains an agent to judge its own attempts, propose targeted revisions, and sustain productive improvement over many rounds.

Key Findings
  • Long-horizon reflection as a transferable meta-skill: training on synthetic multi-round trajectories from machine-learning engineering and algorithmic programming produces an agent that transfers to deep-research benchmarks without domain-specific supervision. This implies the reflective strategy (identify weakness → propose focused change → verify) generalizes across problem types.
  • Complementary capabilities drive test-time gains: reflection raises per-round improvement, while long-horizon execution preserves productivity across rounds. Empirically AREX-2 (built on Qwen3.8-27B) achieves 81.8 on MLE-bench Lite, 70.7 on Frontier-CS, and improves on BrowseComp, HLE, GAIA, and DeepSearchQA as its round budget grows.
  • Training with explicit multi-round trajectories matters: lessons learned are not recoverable from finished solutions alone—supervising the iterative process (what to change after feedback) yields sustained gains that scale with more rounds.
Who it's for and tradeoffs

Great fit if you care about agents that can autonomously refine complex artifacts (code, training pipelines, research answers) across many iterations and want a model that improves given extra compute budget. AREX-2 is valuable for researchers building self-improving agents, benchmark designers studying iteration dynamics, and teams exploring automated experiment loops.

Look elsewhere if you need a drop-in product for turnkey automation: AREX-2 demonstrates a training recipe and agent behavior rather than an out-of-the-box system; it requires training data of long-horizon trajectories and computational budget to realize multi-round gains. Also, current results reflect training on two supervised domains (ML engineering, algorithmic programming) and a Qwen3.8-27B backbone—generalization to very different modalities or lower-resource settings may need further domain expansion or scaling.

Information

  • Websitearxiv.org
  • OrganizationsBeijing Academy of Artificial Intelligence (BAAI)
  • AuthorsHongjin Qian, Chaofan Li, Kun Luo, Wenqing Wei, Jianlyu Chen, Shuqi Lu, Yuyang Hu, Hongwang Xiao, Hui Wang, Chaozhuo Li …
  • Published date2026/09/29

More Items

Analyzes how proposer–solver loops in self-evolving search agents can develop shared errors (co-cheating) that inflate internal rewards; introduces Multi-Sample Verification and CrossFit (cross-fitted scoring with partitioned sources) to reduce false agreement and improve downstream search performance.

Co-evolves candidate solutions and web-search queries to help LLM-driven evolutionary discovery, using a retrieval gate plus bilevel inner/outer loops that refine queries, rank documents by predicted solution value, and generate evaluated candidates.

Provides a plug-and-play harness that makes existing agents omni-native by exposing hierarchical multimodal Skills, a standardized execution interface, dependency-aware orchestration, and a persistent Asset Registry. Represents multi-asset workflows as Declare Execution Graphs to schedule concurrent operations and enable cross-turn reuse across interchangeable execution backends.