Collecting successful trajectories for complex terminal tasks is essential for post-training coding agents. However, solving difficult tasks often relies on specialized harnesses that incorporate domain assumptions unavailable during real-world deployment; naive SFT on these traces causes verifier leakage and harness over-fitting. Researchers from the University of Maryland and Tencent AI Lab introduce Recursive Self-Rewrite (RSR, arXiv:2610.02826). Using a single foundation model (Qwen-3.8-27B), RSR pairs a planner (extracting operational runbooks), a critic (eliminating leakage), and an executor (replaying runbooks in fresh isolated sandboxes). RSR expands 2,001 specialized source rollouts into 11,094 clean general-harness trajectories, boosting Terminal-Bench 2 Pass@3 from 57.0% to 74.2% and surging Terminal-Bench Hard success from 39.0% to 63.0%.
Key Takeaways
- ✓Pioneers Recursive Self-Rewrite (RSR) to distill specialized harness experiences into clean general-harness SFT trajectories
- ✓Expands 2,001 noisy exploration rollouts into 11,094 fully reproducible trajectories in pristine sandboxes
- ✓Drives Terminal-Bench 2 Pass@3 from 57.0% to 74.2% and lifts Terminal-Bench Hard success from 39.0% to 63.0%

Developer 3s Key Decision Metrics
Turn your technical choice into a development budget
Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
核心背景与行业痛点
Training autonomous terminal agents to resolve complex operating system and codebase failures requires high-quality demonstration trajectories. To resolve difficult long-horizon tasks during data collection, practitioners routinely deploy specialized harnesses that inject auxiliary prompts, debug hooks, or customized toolkits. While specialized harnesses achieve successful runs, training foundation models directly on these traces triggers catastrophic harness overfitting and verifier leakage; when deployed in clean production terminal environments lacking specialized scaffolding, agents falter immediately.
架构亮点与底层机制
Researchers from the University of Maryland and Tencent AI Lab introduce Recursive Self-Rewrite (RSR, arXiv:2610.02826):
- Unified Tripartite Agent Ecosystem: Operates with a single foundation model (Qwen-3.8-27B) configured into distinct roles—Planner, Critic, and Executor—establishing a self-contained data distillation loop without proprietary teacher APIs.
- Procedural Runbook Synthesis: The Planner distills sprawling, harness-entangled execution logs into clean, generalized operational runbooks.
- Leakage Prevention via Recursive Revision: The Critic inspects candidate procedures to eliminate test verifier leakage, solution shortcuts, and environment-dependent artifacts, orchestrating multi-round recursive rewrites.
- Clean-Sandbox Verification: The Executor attempts the validated runbook in a fresh, uninstrumented baseline terminal sandbox, admitting only those traces that prove 100% reproducible in bare environments.
权威 Benchmark 与实测跑分对比
Benchmarked across 3,000 diverse Linux system tasks and official Terminal-Bench suites:
- 5.5x Multiplier of Verified Data: RSR transforms 2,001 raw source rollouts across heterogeneous harnesses into 11,094 clean, general-harness trajectories.
- Terminal-Bench 2 Pass@3 Jumps to 74.2%: Post-training on RSR data elevates Pass@3 from 57.0% to 74.2% on Terminal-Bench 2, and pushes Terminal-Bench 4 from 1.5% to 9.1%.
- Catapults Terminal-Bench Hard to 63.0%: Achieves a massive leap from 39.0% to 63.0% on Terminal-Bench Hard, while lifting Long-Horizon Process Rewards from 0.21 to 0.29.
开发者实战落地与开箱指南
RSR provides a production-tested data synthesis pipeline for engineering teams building DevOps and software engineering copilots. By converting tool-heavy exploration rollouts into clean, general-purpose terminal execution traces, engineering teams can post-train foundation models to navigate standard bash environments without reliance on auxiliary runtime interventions.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.