AIAny
Icon for item

Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge

Teaches small reasoning models to diagnose when additional internal thinking is insufficient and to selectively query stronger models; introduces FlyBy, a supervised + cost-aware RL framework that learns when/what to ask. Improves pass@ metrics on hard benchmarks while reducing serving cost.

Introduction

Current test-time scaling via longer chains of thought often assumes more internal computation always helps. This paper shows that for small reasoning models additional self-refinement mostly consolidates already-reachable solutions instead of unlocking new ones, exposing two distinct failure modes—execution bottlenecks (fixable by reflection) and knowledge bottlenecks (requiring external parametric knowledge). Motivated by this, the authors propose FlyBy: a selective querying pipeline that lets small models reason, diagnose unresolved gaps, and query stronger models only when a knowledge bottleneck is detected.

Key Findings
  • Distinguishing failure modes: empirical interventions reveal that many errors stem from either execution mistakes (solvable by short refinement) or genuine knowledge gaps (requiring external models), which calls for different remedies rather than uniformly more thinking.
  • FlyBy framework: combines supervised fine-tuning to bootstrap multi-depth query actions with cost-aware reinforcement learning to decide whether to query, what to ask, and how much compute to spend, prioritizing cost–accuracy tradeoffs.
  • Empirical gains: on 1,158 hard problems across six benchmarks, FlyBy-4B improves pass@8 and pass@1 relative to larger baselines at substantially lower serving cost; scaling to FlyBy-8B gives further gains.
Who it's for and tradeoffs

Great fit if you deploy small-to-medium reasoning models and need a practical way to extend factual reach without always serving large models — FlyBy reduces average serving cost by querying only when necessary. Look elsewhere if you cannot tolerate any added system complexity (multi-depth querying, supervised labels for query decisions, or cost-aware RL) or if strict latency constraints forbid occasional external queries.

Information

  • Websitearxiv.org
  • OrganizationsAffiliation: KAIST, Affiliation: DeepAuto.ai(†\dagger: Equal advising)
  • AuthorsChanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
  • Published date2026/09/28

More Items

Shows that private post-training changes leave measurable "behavioral shadows" in task-unrelated single-word outputs and proposes Active Taskless Distillation (ATD) to transfer capabilities to a student using only those single-word teacher responses, without teacher logits or parameters.

Improves test-time scaling of looped transformers by adaptively assigning extra recurrent iterations to tokens that benefit most. TaH2 is a post-training method that jointly trains an iteration decider with the backbone using lookahead depth supervision, boosting the accuracy–compute slope and peak accuracy on challenging benchmarks.

Normalizes each domain's teacher feedback spread during multi-teacher on-policy distillation so that no domain (e.g., instruction following) overwhelms others, improving student recovery of specialist skills and raising average scores across benchmarks.