AIAny
Icon for item

Questioning the Questions: Sustaining Self-Evolution in Reasoning Models

Analyzes why self-evolving reasoning models collapse under repeated self-training and proposes R-Quest: a feedback-driven pipeline that trains solvers to reject invalid questions and uses a frozen base model to detect task-level repetition, filtering training data to sustain multi-round gains.

Introduction

Most self-play/self-evolution pipelines assume generated training questions remain helpful across rounds — but they often do not. The paper shows that two overlooked failure modes — rising rates of ill-posed (invalid) questions and repeated mathematical task structures expressed differently — cause late-stage collapse in solver performance. R-Quest addresses both by adding validity and novelty feedback into the questioner–solver loop, producing sustained improvements over many rounds.

Key Findings
  • Invalid questions accumulate over rounds and answer-consistency filtering can amplify them; training a solver to explicitly recognize and reject invalid prompts reduces this noise and improves downstream learning. This means self-generated corpora need an internal validity signal, not just answer agreement.
  • Surface-level lexical diversity measures miss mathematically equivalent tasks; using a frozen base model to compare sampled question pairs detects task-level repetition even when wording changes, preventing diversity collapse and over-concentration on a few task types.
  • Combining validity and novelty feedback (R-Quest) yields stable gains across ten self-evolution rounds, outperforming R-Zero by a large margin (example: +17.32 points on Qwen3-4B-Base at round ten) and keeping checkpoints consistently above the base model.
Who it's for and tradeoffs

Great fit if you research or build self-training/self-play loops for LLM reasoning, curriculum generation, or automated dataset bootstrapping and need methods to sustain multi-round improvement. The approach is practical: both feedback signals are generated with the evolving models (no continual calls to stronger external models), and it targets common quality failures in generated question pools.

Look elsewhere if your primary issue is model capacity or data-domain mismatch rather than noisy/redundant self-generated prompts; R-Quest mitigates poor training signal quality but does not address fundamental model architecture limits or out-of-domain generalization.

Method overview

R-Quest alternates questioner and solver updates while (1) training the solver to flag and reject invalid questions, then using those validity judgments to penalize invalid question generation, and (2) using a frozen base model to perform sampled pairwise comparisons that produce novelty feedback at the task level. The pipeline restricts questioner difficulty rewards to items judged both valid and novel, and filters solver training data accordingly — a lightweight control that targets sustained improvement rather than short-lived gains.

Information

  • Websitearxiv.org
  • OrganizationsWashington University in St. Louis, University of Michigan, Ann Arbor
  • AuthorsJinyuan Li, Chengsong Huang, Langlin Huang, Donghong Cai, Shiping Gao, Yuyi Yang, Jiaxin Huang
  • Published date2026/10/03

More Items

Compresses recurrent states in linear-attention LLMs via spatial–temporal post-training quantization, allocating bits by error lifetime and per-row impact to preserve accuracy (6-bit ≈ FP32) while cutting serving memory up to 68.7%.

Proposes TRACE, an FP4 quantization framework for RL of MoE LLMs that uses rollout-side FP4 outcomes to guide training-side rounding and caches mantissa/scale from deeper layers to limit overhead. Preserves BF16-level RL performance while enabling up to 5.4× rollout speedup.

Analyzes cross-tokenizer on-policy distillation (OPD) and shows that focusing supervision on strictly aligned token positions—or a student-selected top-16 subset of the shared vocabulary—retains most distillation gains, while adding span-level MSE supervision reduces downstream accuracy.