AIAny
Icon for item

EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery

Co-evolves candidate solutions and web-search queries to help LLM-driven evolutionary discovery, using a retrieval gate plus bilevel inner/outer loops that refine queries, rank documents by predicted solution value, and generate evaluated candidates.

Introduction

Most evolutionary search with fixed LLMs stalls when external knowledge is needed but retrieval keeps returning stale pages. EvoDuet flips this interaction: instead of blindly adding web search, it co-evolves solutions and queries so the agent only fetches new evidence when it predicts a knowledge gap, and uses evaluated outcomes to guide future retrieval.

Key Findings
  • Bilevel co-evolution: an inner loop refines search queries and ranks verified documents by the solution scores they are predicted to yield, while an outer loop generates candidate solutions in parallel from those documents and records real evaluation outcomes for later retrieval.
  • Retrieval gating: the LLM assesses whether it needs external documents, can reuse stored evidence, or should proceed without retrieval, reducing redundant fetches as solutions evolve.
  • Empirical gains: with one candidate per iteration across 21 optimization tasks, EvoDuet raised normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash; Qwen3.5-9B showed no benefit. Best runs surpassed previous bests on eight tasks and matched three more.
Who it's for and trade-offs

Great fit if you run LLM-guided evolutionary or black-box optimization and need better use of web evidence—especially when models lack up-to-date or domain-specific knowledge. Look elsewhere if your workflow already uses highly specialized retrieval pipelines, if retrieval latency dominates cost, or if you only run tiny models that cannot leverage retrieved context; EvoDuet’s benefits depend on model capability and the cost of extra retrieval/evaluation steps.

Methodological note

EvoDuet treats retrieval as part of the optimization state rather than an external oracle: ranked, evidence-grounded documents become inputs to parallel candidate generators, and evaluated outcomes are fed back to shape future query construction. This looped, data-driven coupling between search and solution generation is the core mechanism that prevents repetitive retrieval and aligns fetched evidence with actual improvement in objective scores.

Information

  • Websitearxiv.org
  • OrganizationsUniversity of Minnesota, KAIST, Seoul National University, Hanyang University
  • AuthorsYoung-Jun Lee, Jinheon Baek, Soyeong Jeong, Minki Kang, Seungyeon Jwa, Jonghyun Choi, Seungho Han, Dongyeop Kang
  • Published date2026/09/30

Categories

More Items

Analyzes how proposer–solver loops in self-evolving search agents can develop shared errors (co-cheating) that inflate internal rewards; introduces Multi-Sample Verification and CrossFit (cross-fitted scoring with partitioned sources) to reduce false agreement and improve downstream search performance.

Defines and evaluates AREX-2, an LLM agent that iteratively self-improves at test time via reflection and long-horizon execution. Trained on long-horizon improvement trajectories from ML engineering and algorithmic programming (built on Qwen3.8-27B), it scales with more rounds and achieves strong benchmark scores.

Provides a plug-and-play harness that makes existing agents omni-native by exposing hierarchical multimodal Skills, a standardized execution interface, dependency-aware orchestration, and a persistent Asset Registry. Represents multi-asset workflows as Declare Execution Graphs to schedule concurrent operations and enable cross-turn reuse across interchangeable execution backends.