Developer 3s Key Decision Metrics
Terminal agents operate through stochastic model generations, but a single errant command can irreversibly degrade sandbox environments. Mid-Harness introduces test-time compute scaling at the boundary between model and harness. By sampling and verifying candidate actions before terminal dispatch without modifying the underlying generator or harness, Mid-Harness boosts Pass@1 on TerminalBench-Lite from 50.00% to 68.03% under a GPT-5.6 Sol verifier, outperforming whole-trajectory scaling at lower token expenditure.
Key Takeaways
- ✓Allocates test-time compute at model-harness boundary, lifting TerminalBench-Lite Pass@1 from 50.00% to 68.03%
- ✓Pairwise action verification and distillation enable compact local models to filter errant commands effectively
- ✓Reduces estimated token overhead by over 35% compared to whole-trajectory sampling while increasing task completion
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.