Hard improvements in research and engineering rarely come from a single pass; they come from repeated try-measure-revise cycles. AREX-2 moves that loop inside the model: instead of asking the harness to orchestrate retries, it trains an agent to judge its own attempts, propose targeted revisions, and sustain productive improvement over many rounds.
Key Findings
- Long-horizon reflection as a transferable meta-skill: training on synthetic multi-round trajectories from machine-learning engineering and algorithmic programming produces an agent that transfers to deep-research benchmarks without domain-specific supervision. This implies the reflective strategy (identify weakness → propose focused change → verify) generalizes across problem types.
- Complementary capabilities drive test-time gains: reflection raises per-round improvement, while long-horizon execution preserves productivity across rounds. Empirically AREX-2 (built on Qwen3.8-27B) achieves 81.8 on MLE-bench Lite, 70.7 on Frontier-CS, and improves on BrowseComp, HLE, GAIA, and DeepSearchQA as its round budget grows.
- Training with explicit multi-round trajectories matters: lessons learned are not recoverable from finished solutions alone—supervising the iterative process (what to change after feedback) yields sustained gains that scale with more rounds.
Who it's for and tradeoffs
Great fit if you care about agents that can autonomously refine complex artifacts (code, training pipelines, research answers) across many iterations and want a model that improves given extra compute budget. AREX-2 is valuable for researchers building self-improving agents, benchmark designers studying iteration dynamics, and teams exploring automated experiment loops.
Look elsewhere if you need a drop-in product for turnkey automation: AREX-2 demonstrates a training recipe and agent behavior rather than an out-of-the-box system; it requires training data of long-horizon trajectories and computational budget to realize multi-round gains. Also, current results reflect training on two supervised domains (ML engineering, algorithmic programming) and a Qwen3.8-27B backbone—generalization to very different modalities or lower-resource settings may need further domain expansion or scaling.