AIAny
Icon for item

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Provides a benchmark and protocol to evaluate agents that iteratively edit executable policies under a fixed interaction budget, recording full execution–feedback–revise trajectories. Built from compact RL environments with trajectory-level diagnostics and hidden held-out validation.

Introduction

Most benchmarks measure only final task scores; they hide the information-acquisition and revision process that matters when an agent must evolve executable policies under limited feedback. EvoPolicyGym flips that frame: it treats the edit–submit–feedback loop itself as the evaluated object, forcing agents to decide what to probe, when to explore, and how to convert sparse rollout evidence into robust code changes.

Key Findings
  • Formalizes Autonomous Policy Evolution as a controlled evaluation setting so what: makes policy-search behavior (budget allocation, checkpoint selection, edit structure) measurable rather than an implementation artifact.
  • Instantiates a Core16 suite of compact RL environments so what: enables cross-task comparison under a common 128-episode interaction budget and hidden held-out selection.
  • Trajectory-level diagnostics reveal mechanism differences so what: strong agents do more than win isolated tasks—they discover task-appropriate mechanisms, preserve promising candidates, and balance structural synthesis vs parameter tuning.
  • Baseline leaderboard outcomes (e.g., top aggregate rank for a strong LLM-based agent) so what: highlight coverage and consistency across environments as distinct from isolated first-place wins.
Who it's for and tradeoffs

Great fit if you want to measure how coding agents transform rollout feedback into concrete policy edits, compare harness–model workflows, or study budget-conditioned search strategies. Look elsewhere if your goal is large-scale robotics deployment, long-horizon simulator-heavy training, or pure end-to-end RL benchmarks—the suite focuses on compact, sandboxed environments and constrains interaction budget by design.

Where it fits

EvoPolicyGym sits between conventional RL leaderboards and open-ended engineering benchmarks: unlike single-score evaluations it records the full revision trajectory; unlike open-ended engineering suites it enforces strict visibility boundaries and a fixed interaction budget so that the evaluation isolates autonomous policy evolution rather than incremental engineering effort.

Information

  • Websitearxiv.org
  • OrganizationsUniversity of Science and Technology of China, The Chinese University of Hong Kong, University of Macau, Tsinghua University, Zhejiang University, Soochow University, Brown University, Shanghai Jiao Tong University
  • AuthorsZhilin Wang, Han Song, Runzhe Zhan, Jusen Du, Jiacheng Chen, Tianle Li, Qingyu Yin, Yulun Wu, Zhennan Shen, Tong Zhu …
  • Published date2026/07/02

More Items

Provides a plug-and-play harness that makes existing agents omni-native by exposing hierarchical multimodal Skills, a standardized execution interface, dependency-aware orchestration, and a persistent Asset Registry. Represents multi-asset workflows as Declare Execution Graphs to schedule concurrent operations and enable cross-turn reuse across interchangeable execution backends.

Introduces an open-source multi-agent harness that automatically constructs, composes, and evolves modular model–harness units for long‑horizon, cross‑domain workflows. Key ingredients include a Host Agent for task decomposition and orchestration, EverOS for durable memory, and a Skill Forge of reusable procedures to improve task coverage via composition.

Trains a single unified multimodal model with reinforcement learning to perform end-to-end self-reflection and iterative image repair — jointly learning the diagnostic (textual) reflection and the flow-based image revisions so credit propagates across rounds without an external verifier.