Terminal coding agents are expanding from standard software engineering into scientific domains like bioinformatics, materials science, and physics. However, building training sandboxes requires executable reference behavior and rigorous verifiers capable of distinguishing true semantic correctness from superficial syntactic mimicry—a process that currently requires tedious manual engineering per task. Researchers from Tsinghua University and Tencent introduce Software-in-the-Loop Reconstruction (SWR, arXiv:2610.02710), a self-supervised environment scaling framework. SWR extracts ground-truth outputs and verifiers from existing executable scientific software workflows by perturbing input configurations into public observations and hidden evaluations. Agents synthesize complete, editable programs from input-output schemas without viewing source code, evaluated by hierarchical verifiers testing semantic fidelity, structural integrity, and anti-shortcut compliance. Implemented across 500 workflows and 46 software suites spanning six scientific domains, SWR generates 1,422 verified trajectories. Supervised fine-tuning of Qwen3.8-27B elevates Terminal-Bench 2 average success from 47.94% to 53.56%, establishing a scalable paradigm for scientific terminal agent training.
Key Takeaways
- ✓Introduces SWR to bootstrap self-supervised terminal agent training sandboxes from existing scientific software workflows
- ✓Decouples public observations from hidden evaluations, using hierarchical verifiers to block shortcuts during program reconstruction
- ✓Drives Qwen3.8-27B performance on Terminal-Bench 2 from 47.94% to 53.56%, eclipsing all token-matched baseline controls

Developer 3s Key Decision Metrics
Turn your technical choice into a development budget
Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
核心背景与行业痛点
Terminal agents (such as SWE-bench and Terminal-Bench systems) possess remarkable autonomy across Linux bash shells, DevOps scripting, and repository management. However, deploying agents into sophisticated scientific workflows—such as structural biology, computational fluid dynamics, and materials science—encounters a chronic bottleneck: high-fidelity sandbox environments with deterministic verifiers are severely lacking. Manually authoring specialized tasks, reference outputs, and semantic rubrics requires prohibitive engineering overhead, limiting scientific post-training data diversity.
架构亮点与底层机制
Researchers from Tsinghua University and Tencent present Software-in-the-Loop Reconstruction (SWR, arXiv:2610.02710), a self-supervised environment synthesis paradigm:
- Bootstrapping From Existing Executable Workflows: Bypasses manual task creation by tapping into existing mature software workflows across production science domains. By running multiple input configurations, SWR partitions output instances into public demonstration samples and private hidden evaluation sets.
- Black-Box Program Reconstruction: Given natural language goals, input schemas, and public observations, an agent must construct an editable, standalone program inside a clean terminal without seeing the underlying proprietary pipeline source code.
- Hierarchical Anti-Shortcut Verifier: Executes the synthesized script across hidden configurations. The verifier validates floating-point semantic accuracy, structural schema compliance, and inspects code execution to eliminate hardcoded lookup shortcuts.
权威 Benchmark 与实测跑分对比
Instantiated across 500 realistic scientific workflows and 46 distinct software families spanning six major STEM disciplines:
- 1,422 Verified Scientific Trajectories: Within 3 attempts per task, Qwen3.8-Max solves 838 arduous computational tasks, generating 1,422 rigorously verified execution traces (expanded to 3,000 reconstruction examples).
- Terminal-Bench 2 Surges to 53.56%: Fine-tuning Qwen3.8-27B exclusively on SWR data elevates average accuracy on Terminal-Bench 2 from 47.94% to 53.56% across three seeds (+5.62 percentage points).
- Dominates Token-Matched Baseline Controls: Outperforms all four token-matched synthetic corpus control baselines across every single reported evaluation dimension, verifying the cross-domain transferability of software-in-the-loop training.
开发者实战落地与开箱指南
The SWR synthesis framework, verifier rules, and software harnesses are documented on arXiv (arXiv:2610.02710). AI for Science organizations and enterprise DevOps engineering teams can plug existing internal CLI tools and computational workflows into SWR to generate vast pools of high-fidelity terminal training environments without manual annotation, fueling self-evolving post-training for scientific agents.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.