Collecting successful trajectories for complex terminal tasks is essential for post-training coding agents. However, solving difficult tasks often relies on specialized harnesses that incorporate domain assumptions unavailable during real-world deployment; naive SFT on these traces causes verifier leakage and harness over-fitting. Researchers from the University of Maryland and Tencent AI Lab introduce Recursive Self-Rewrite (RSR, arXiv:2610.02826). Using a single foundation model (Qwen-3.8-27B), RSR pairs a planner (extracting operational runbooks), a critic (eliminating leakage), and an executor (replaying runbooks in fresh isolated sandboxes). RSR expands 2,001 specialized source rollouts into 11,094 clean general-harness trajectories, boosting Terminal-Bench 2 Pass@3 from 57.0% to 74.2% and surging Terminal-Bench Hard success from 39.0% to 63.0%.

Key Takeaways

  • ✓Pioneers Recursive Self-Rewrite (RSR) to distill specialized harness experiences into clean general-harness SFT trajectories
  • ✓Expands 2,001 noisy exploration rollouts into 11,094 fully reproducible trajectories in pristine sandboxes
  • ✓Drives Terminal-Bench 2 Pass@3 from 57.0% to 74.2% and lifts Terminal-Bench Hard success from 39.0% to 63.0%
Recursive Self-Rewrite (RSR) Scales Complex Terminal Trajectories: Distilling Specialized Harnesses into General SFT Capabilities
🖼️Official Media
Click to view high-res
🧭

Turn your technical choice into a development budget

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

核心背景与行业痛点

Training autonomous terminal agents to resolve complex operating system and codebase failures requires high-quality demonstration trajectories. To resolve difficult long-horizon tasks during data collection, practitioners routinely deploy specialized harnesses that inject auxiliary prompts, debug hooks, or customized toolkits. While specialized harnesses achieve successful runs, training foundation models directly on these traces triggers catastrophic harness overfitting and verifier leakage; when deployed in clean production terminal environments lacking specialized scaffolding, agents falter immediately.

架构亮点与底层机制

Researchers from the University of Maryland and Tencent AI Lab introduce Recursive Self-Rewrite (RSR, arXiv:2610.02826):

  1. Unified Tripartite Agent Ecosystem: Operates with a single foundation model (Qwen-3.8-27B) configured into distinct roles—Planner, Critic, and Executor—establishing a self-contained data distillation loop without proprietary teacher APIs.
  2. Procedural Runbook Synthesis: The Planner distills sprawling, harness-entangled execution logs into clean, generalized operational runbooks.
  3. Leakage Prevention via Recursive Revision: The Critic inspects candidate procedures to eliminate test verifier leakage, solution shortcuts, and environment-dependent artifacts, orchestrating multi-round recursive rewrites.
  4. Clean-Sandbox Verification: The Executor attempts the validated runbook in a fresh, uninstrumented baseline terminal sandbox, admitting only those traces that prove 100% reproducible in bare environments.

权威 Benchmark 与实测跑分对比

Benchmarked across 3,000 diverse Linux system tasks and official Terminal-Bench suites:

  1. 5.5x Multiplier of Verified Data: RSR transforms 2,001 raw source rollouts across heterogeneous harnesses into 11,094 clean, general-harness trajectories.
  2. Terminal-Bench 2 Pass@3 Jumps to 74.2%: Post-training on RSR data elevates Pass@3 from 57.0% to 74.2% on Terminal-Bench 2, and pushes Terminal-Bench 4 from 1.5% to 9.1%.
  3. Catapults Terminal-Bench Hard to 63.0%: Achieves a massive leap from 39.0% to 63.0% on Terminal-Bench Hard, while lifting Long-Horizon Process Rewards from 0.21 to 0.29.

开发者实战落地与开箱指南

RSR provides a production-tested data synthesis pipeline for engineering teams building DevOps and software engineering copilots. By converting tool-heavy exploration rollouts into clean, general-purpose terminal execution traces, engineering teams can post-train foundation models to navigate standard bash environments without reliance on auxiliary runtime interventions.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.