Most evolutionary search with fixed LLMs stalls when external knowledge is needed but retrieval keeps returning stale pages. EvoDuet flips this interaction: instead of blindly adding web search, it co-evolves solutions and queries so the agent only fetches new evidence when it predicts a knowledge gap, and uses evaluated outcomes to guide future retrieval.
Key Findings
- Bilevel co-evolution: an inner loop refines search queries and ranks verified documents by the solution scores they are predicted to yield, while an outer loop generates candidate solutions in parallel from those documents and records real evaluation outcomes for later retrieval.
- Retrieval gating: the LLM assesses whether it needs external documents, can reuse stored evidence, or should proceed without retrieval, reducing redundant fetches as solutions evolve.
- Empirical gains: with one candidate per iteration across 21 optimization tasks, EvoDuet raised normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash; Qwen3.5-9B showed no benefit. Best runs surpassed previous bests on eight tasks and matched three more.
Who it's for and trade-offs
Great fit if you run LLM-guided evolutionary or black-box optimization and need better use of web evidence—especially when models lack up-to-date or domain-specific knowledge. Look elsewhere if your workflow already uses highly specialized retrieval pipelines, if retrieval latency dominates cost, or if you only run tiny models that cannot leverage retrieved context; EvoDuet’s benefits depend on model capability and the cost of extra retrieval/evaluation steps.
Methodological note
EvoDuet treats retrieval as part of the optimization state rather than an external oracle: ranked, evidence-grounded documents become inputs to parallel candidate generators, and evaluated outcomes are fed back to shape future query construction. This looped, data-driven coupling between search and solution generation is the core mechanism that prevents repetitive retrieval and aligns fetched evidence with actual improvement in objective scores.