While test-time compute scaling has propelled reasoning capabilities in large models, small reasoning models (sRMs, 4B-8B) frequently fail when scaling thinking tokens: self-refinement merely consolidates probability onto already reachable hypotheses, failing completely when missing parametric knowledge. KAIST researchers introduce FlyBy, a selective querying framework that formally decouples execution bottlenecks (recoverable through internal reflection) from knowledge bottlenecks (requiring external teacher queries). Optimized via cost-aware reinforcement learning, FlyBy teaches small models when to think, when to ask, and how much compute budget to expend. Evaluated across 1,158 hard problems on six benchmarks, FlyBy-4B reaches 45.96% pass@8, outperforming the much larger Qwen3-14B (41.64%) while slashing serving costs by 2.7x (a 63% reduction).

Key Takeaways

  • ✓Diagnosing Blind Thinking in sRMs: Empirical interventions demonstrate that self-refinement primarily concentrates probability mass on pre-reachable states; lacking key facts, extra test-time compute leads to hallucination loops.
  • ✓Execution vs. Knowledge Bottleneck Decoupling: FlyBy introduces formal state interventions to separate execution errors (amenable to reflection) from knowledge bottlenecks (requiring external model invocation).
  • ✓4B Surpasses 14B at 63% Lower Cost: On 1,158 difficult multi-domain problems, FlyBy-4B scores 45.96% pass@8 (beating Qwen3-14B's 41.64%) at 2.7x lower serving cost, while FlyBy-8B scales to 51.81%.
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Core Background & Industry Pain Points The surge of test-time compute scaling and extended Chain-of-Thought (CoT) reasoning has highlighted Small Reasoning Models (sRMs, 4B-8B) as prime candidates for edge deployment, local code assistants, and latency-sensitive agent workflows. However, practitioners frequently encounter the "Blind Thinking Trap": 1. Parametric Capacity Ceilings: Small models contain finite parametric knowledge. When presented with unseen facts, APIs, or domain mechanics, extending internal reasoning does not magically synthesize missing knowledge; instead, it reinforces hallucinated trajectories; 2. Wasted Test-Time Compute: Mobile and edge batteries burn thousands of futile reasoning tokens exploring dead-end branches, inflating latency while failing to reach correct answers. ### Architectural Highlights & Underlying Mechanics KAIST researchers introduced FlyBy, a selective querying framework designed to discipline test-time compute in small reasoning models: 1. State-Intervention Probing: Intervening at intermediate reasoning states empirically bifurcates failures into Execution Bottlenecks (where correct paths are reachable and self-refinement succeeds) and Knowledge Bottlenecks (where essential external information is strictly required); 2. Multi-Depth Query Actions: SFT equips the sRM with metacognitive introspection, prompting it to self-diagnose reasoning impasses and frame concise query prompts targeting unresolved knowledge gaps; 3. Cost-Aware Reinforcement Learning: Calibrates the agent via an RL reward function penalizing unnecessary external queries, optimizing the tradeoff between local token generation and selective teacher invocation. ### Benchmark & Experimental Validation Evaluated across 1,158 challenging problems on six rigorous reasoning benchmarks: - 4B Outperforms 14B Baselines: FlyBy-4B achieves 45.96% pass@8, decisively outperforming standalone Qwen3-14B (41.64%), while also beating Qwen3-8B in greedy pass@1 (16.85% vs 15.31%); - 63% Cost Reduction (2.7x Efficiency): By restricting queries strictly to verified knowledge roadblocks, FlyBy-4B cuts blended inference costs by 2.7x compared to medium-scale monolithic models; - Scaling to 8B: FlyBy-8B elevates pass@8 to 51.81%, validating that selective collaboration scales predictably across model sizes. ### Engineering Takeaways & Practical Guide - Paper Reference: Full details are available in arXiv:2609.34327; - Blueprint for Edge-Cloud Coding Agents: FlyBy offers a blueprint for agent developers: maintain a local 4B/8B sRM for fast terminal loopback, file navigation, and deterministic edits, while dispatching surgical RPC requests to frontier cloud models only when diagnosing knowledge bottlenecks; - RL Alignment Insight: Prompt engineering alone cannot prevent over-confidence; cost-penalized RL is required to align self-evaluation with query efficiency.