Terminal agents operate through stochastic model generations, but a single errant command can irreversibly degrade sandbox environments. Mid-Harness introduces test-time compute scaling at the boundary between model and harness. By sampling and verifying candidate actions before terminal dispatch without modifying the underlying generator or harness, Mid-Harness boosts Pass@1 on TerminalBench-Lite from 50.00% to 68.03% under a GPT-5.6 Sol verifier, outperforming whole-trajectory scaling at lower token expenditure.

Key Takeaways

  • ✓Allocates test-time compute at model-harness boundary, lifting TerminalBench-Lite Pass@1 from 50.00% to 68.03%
  • ✓Pairwise action verification and distillation enable compact local models to filter errant commands effectively
  • ✓Reduces estimated token overhead by over 35% compared to whole-trajectory sampling while increasing task completion
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点 Terminal coding agents increasingly interact with active shell environments. However, autoregressive models produce stochastic commands where a single destructive command (e.g., misconfigured pip install, accidental directory overwrite) irreparably poisons the execution container, causing downstream tasks to fail even if the model possessed viable alternatives. Conventional trajectory-level sampling incurs massive token costs without preventing early deterministic errors. ### 架构亮点与底层机制 Mid-Harness allocates test-time compute directly at the model-harness boundary without modifying base weights or execution harnesses. Before dispatching any command to the shell, it samples multiple action candidates and subjects them to targeted verification. Under symmetric model configurations, pairwise verification isolates optimal actions with high signal-to-noise ratios. Furthermore, distilling strong verifier feedback into compact models enables local, low-latency pre-execution gating. ### 权威 Benchmark 与实测跑分对比 Evaluated on the authoritative TerminalBench-Lite environment: 1. Pass@1 Surges by 18%: Using TMAX-9B as generator, Mid-Harness with an 8-action GPT-5.6 Sol verifier elevates Pass@1 from 50.00% to 68.03% (+18.03 percentage points). 2. Superiority of Pairwise Verification: When self-verifying with TMAX-9B, pairwise evaluation outperforms point-wise scoring, matching distilled enterprise benchmarks. 3. 35%+ Token Savings Over Trajectory Scaling: Combining action scaling at critical branch points achieves higher task completion with >35% fewer tokens than repeatedly sampling full trajectories. ### 开发者实战落地与开箱指南 Mid-Harness provides a modular blueprint for terminal assistant builders. Integrating a verification interceptor at the MCP tool call or Bash execution gateway prevents container contamination and drastically enhances stability across long-horizon software engineering workloads.