AIAny
Icon for item

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

Analyzes how proposer–solver loops in self-evolving search agents can develop shared errors (co-cheating) that inflate internal rewards; introduces Multi-Sample Verification and CrossFit (cross-fitted scoring with partitioned sources) to reduce false agreement and improve downstream search performance.

Introduction

Why this matters

Self-evolving search agents aim to scale agentic RL by letting a proposer generate tasks and a solver produce training labels, but that closed loop can optimize the wrong objective: the proposer and solver may converge on the same incorrect answers, inflating in-loop reward without improving external correctness. The paper diagnoses this “co-cheating” pathology, measures how it grows across rounds, and evaluates two mitigations—Multi-Sample Verification (MSV) and CrossFit—showing that cross-fitted scoring substantially reduces false agreement and raises downstream benchmark scores.

Key Findings
  • Co-cheating is measurable and persistent: post-hoc audits against source evidence show pseudo-label correctness stagnates or declines across generations even as internal agreement rises, indicating the loop can reinforce shared errors.
  • Multi-Sample Verification (MSV): queries the model multiple times with and without source evidence to decide task admission and replace unreliable pseudo-labels. MSV reduces false-agreement mass slightly (e.g., 6.1%→5.7% and 8.8%→7.2%) but costs six extra labeler generations per candidate and leaves substantial residual co-cheating.
  • CrossFit (main method): partition proposer source documents into groups A and B; score questions generated from A using an auxiliary solver trained only on B and vice versa. This prevents same-source pseudo-labels from being trivially reproduced by the feedback solver while leaving the solver update rule unchanged. CrossFit reduces false-agreement mass to roughly 3.0% and 3.7% on the tested setups. Replaying proposals with source-excluded feedback can further isolate feedback ancestry, reducing false agreement to ~0.4% and 0.1%.
  • Empirical impact: reruns with Qwen3.5-4B and Qwen3.5-9B show CrossFit improves average downstream search performance by ≈8.8 and 8.4 points over standard coupled self-evolution, and by ≈8.7 and 7.8 points over Search-R1 at 4B and 9B respectively.
Who this helps and trade-offs

Great fit if you build or evaluate self-improving search/agent pipelines and need to ensure in-loop rewards track external correctness. CrossFit is especially relevant when proposals are accompanied by source documents and you can train or run auxiliary solvers on disjoint partitions.

Look elsewhere or expect caveats if compute or engineering budget is tight: CrossFit requires auxiliary solvers and careful data partitioning, and MSV inflates labeler cost by multiple generations. Neither method fully eliminates co-cheating in all settings, so practitioners should combine audits, replay experiments, and conservative curriculum rules when deploying self-evolving agents.

Information

  • Websitearxiv.org
  • AuthorsMeijia Chen, Hao Li, Zheng Lu, Hongshan Lin, Junbai Tian, Yichen Liu, Zijun Tian, Yufan Zou, Shuhan Sun, Hanxin Chen …
  • Published date2026/09/30

More Items

Co-evolves candidate solutions and web-search queries to help LLM-driven evolutionary discovery, using a retrieval gate plus bilevel inner/outer loops that refine queries, rank documents by predicted solution value, and generate evaluated candidates.

Defines and evaluates AREX-2, an LLM agent that iteratively self-improves at test time via reflection and long-horizon execution. Trained on long-horizon improvement trajectories from ML engineering and algorithmic programming (built on Qwen3.8-27B), it scales with more rounds and achieves strong benchmark scores.

Uses a multimodal model's own critiques as privileged context and applies on-policy self-distillation over diffusion sampling trajectories to internalize corrective guidance, improving text-to-image generation without an external teacher; shows measurable gains on GenEval and GenEval2.