Recent attempts to scale tool-use post-training for autonomous agents have focused primarily on synthesizing standalone execution environments. However, environments constitute merely one component of a holistic agentic interaction system comprising environments, tasks, harnesses, and evaluators; scaling environments in isolation yields inconsistent training signals. Researchers from Fudan University introduce WEFT (arXiv:2609.36887), a whole-system evolution framework for tool-use post-training. WEFT coordinates environment breadth, task complexity, and interaction diversity, applying execution-driven self-evolution to attribute failures and iteratively revise flawed components. To ensure optimization stability and concurrency reliability, WEFT introduces prefix-preserving sampling, atomic-turn credit assignment, and MegaMCP—a Model Context Protocol infrastructure providing isolated, recoverable state management across concurrent rollouts. WEFT-8B and WEFT-14B outperform matched environment-scaling baselines across BFCL V4, tau^2-Bench, and Claw-Eval, with WEFT-14B beating Agent-World-14B by 6.41, 2.23, and 12.27 percentage points, while WEFT-35B-A3B excels on long-horizon benchmarks like Toolathlon-Verified.
Key Takeaways
- ✓Pioneers WEFT, co-evolving environments, tasks, harnesses, and evaluators to overcome the limits of isolated environment synthesis
- ✓Integrates prefix-preserving sampling, atomic-turn credit assignment, and MegaMCP isolation for resilient tool-use post-training
- ✓WEFT-14B surges past Agent-World-14B by 6.41 pp on BFCL V4, 2.23 pp on tau^2-Bench, and 12.27 pp on Claw-Eval

Developer 3s Key Decision Metrics
Turn your technical choice into a development budget
Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
核心背景与行业痛点
Scaling tool-use capabilities through post-training is foundational for modern autonomous coding and enterprise agents. However, recent paradigms primarily emphasize synthesizing executable sandbox environments (such as Agent-World). In practice, an environment is only one constituent of an agentic interaction system that equally depends on nuanced task definitions, execution harnesses, and robust evaluators. Scaling environments in isolation produces severe cross-component incoherence, generating noisy reward signals that destabilize policy optimization and trigger overfitting to synthetic quirks.
架构亮点与底层机制
Researchers from Fudan University present WEFT (Whole-system Evolution For Tool-use Post-training, arXiv:2609.36887):
- Whole-System Co-Construction: Jointly scales the entire agentic topology across environmental diversity (spanning diverse enterprise APIs), task complexity, and multi-turn interaction modalities.
- Execution-Driven Self-Evolution: Inspects full execution traces and runtime state invariants to pinpoint root causes behind failures across environments, harnesses, or rubrics, applying targeted updates verified by follow-up rollouts.
- Stable Post-Training Optimizations: Introduces prefix-preserving sampling to safeguard confirmed progress during RL exploration, coupled with atomic-turn credit assignment to concentrate policy gradients on critical tool-invocation decisions.
- MegaMCP Infrastructure: Integrates with the Model Context Protocol (MCP) standard to orchestrate isolated, snapshot-recoverable state management across massive concurrent agent rollouts over shared backend services.
权威 Benchmark 与实测跑分对比
Empirically evaluated against state-of-the-art environment-scaling baselines:
- Dominates Across BFCL V4, tau^2-Bench, and Claw-Eval: Both WEFT-8B and WEFT-14B establish new performance highs against identically sized foundation models.
- Up to +12.27 Percentage Point Gains Over Agent-World-14B: WEFT-14B outperforms Agent-World-14B by 6.41 percentage points on BFCL V4, 2.23 points on tau^2-Bench, and an extraordinary 12.27 percentage points on Claw-Eval.
- Robust Long-Horizon Scaling: The scaled WEFT-35B-A3B variant maintains superior accuracy on arduous long-horizon task benchmarks, including Toolathlon-Verified and AutomationBench.
开发者实战落地与开箱指南
WEFT provides a definitive blueprint for machine learning engineers building tool-using agents. By replacing isolated environment generation with whole-system evolution and deploying MegaMCP for resilient state management, teams can construct stable, sample-efficient reinforcement learning pipelines that transfer reliably to production APIs and complex multi-agent workflows.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.