Finds token- and API-cost-saving harness mechanisms for long-horizon coding agents using automated recursive self-improvement; packages four surviving mechanisms (action fusion, context compaction, observation archiving, delegated reading) to cut recorded token traffic ~44.7–49.0% and API cost by about one third while preserving most capability.
Measures how individual harness components—planning, action space, and context management—affect coding agents' success, cost, and behavior. Uses a modular harness across 176 matched settings on SWE-Bench Verified and Terminal-Bench 2.1 to isolate component effects and surface model- and budget-dependent trade-offs.
Provides a domain-agnostic world-modeling framework that factorizes latent targets into orthogonal predictive components, with dedicated prediction branches and synthesis for multi-domain forecasting and intervention. Demonstrates improved dynamics and long-horizon rollouts across seven domains and includes experimental biological validation.