Most self-play/self-evolution pipelines assume generated training questions remain helpful across rounds — but they often do not. The paper shows that two overlooked failure modes — rising rates of ill-posed (invalid) questions and repeated mathematical task structures expressed differently — cause late-stage collapse in solver performance. R-Quest addresses both by adding validity and novelty feedback into the questioner–solver loop, producing sustained improvements over many rounds.
Key Findings
- Invalid questions accumulate over rounds and answer-consistency filtering can amplify them; training a solver to explicitly recognize and reject invalid prompts reduces this noise and improves downstream learning. This means self-generated corpora need an internal validity signal, not just answer agreement.
- Surface-level lexical diversity measures miss mathematically equivalent tasks; using a frozen base model to compare sampled question pairs detects task-level repetition even when wording changes, preventing diversity collapse and over-concentration on a few task types.
- Combining validity and novelty feedback (R-Quest) yields stable gains across ten self-evolution rounds, outperforming R-Zero by a large margin (example: +17.32 points on Qwen3-4B-Base at round ten) and keeping checkpoints consistently above the base model.
Who it's for and tradeoffs
Great fit if you research or build self-training/self-play loops for LLM reasoning, curriculum generation, or automated dataset bootstrapping and need methods to sustain multi-round improvement. The approach is practical: both feedback signals are generated with the evolving models (no continual calls to stronger external models), and it targets common quality failures in generated question pools.
Look elsewhere if your primary issue is model capacity or data-domain mismatch rather than noisy/redundant self-generated prompts; R-Quest mitigates poor training signal quality but does not address fundamental model architecture limits or out-of-domain generalization.
Method overview
R-Quest alternates questioner and solver updates while (1) training the solver to flag and reject invalid questions, then using those validity judgments to penalize invalid question generation, and (2) using a frozen base model to perform sampled pairwise comparisons that produce novelty feedback at the task level. The pipeline restricts questioner difficulty rewards to items judged both valid and novel, and filters solver training data accordingly — a lightweight control that targets sustained improvement rather than short-lived gains.