Current test-time scaling via longer chains of thought often assumes more internal computation always helps. This paper shows that for small reasoning models additional self-refinement mostly consolidates already-reachable solutions instead of unlocking new ones, exposing two distinct failure modes—execution bottlenecks (fixable by reflection) and knowledge bottlenecks (requiring external parametric knowledge). Motivated by this, the authors propose FlyBy: a selective querying pipeline that lets small models reason, diagnose unresolved gaps, and query stronger models only when a knowledge bottleneck is detected.
Key Findings
- Distinguishing failure modes: empirical interventions reveal that many errors stem from either execution mistakes (solvable by short refinement) or genuine knowledge gaps (requiring external models), which calls for different remedies rather than uniformly more thinking.
- FlyBy framework: combines supervised fine-tuning to bootstrap multi-depth query actions with cost-aware reinforcement learning to decide whether to query, what to ask, and how much compute to spend, prioritizing cost–accuracy tradeoffs.
- Empirical gains: on 1,158 hard problems across six benchmarks, FlyBy-4B improves pass@8 and pass@1 relative to larger baselines at substantially lower serving cost; scaling to FlyBy-8B gives further gains.
Who it's for and tradeoffs
Great fit if you deploy small-to-medium reasoning models and need a practical way to extend factual reach without always serving large models — FlyBy reduces average serving cost by querying only when necessary. Look elsewhere if you cannot tolerate any added system complexity (multi-depth querying, supervised labels for query decisions, or cost-aware RL) or if strict latency constraints forbid occasional external queries.