Addressing the fundamental limitation of benchmarks like SWE-bench that assess coding agents strictly within isolated repositories, researchers from Zhejiang University introduced WideSWE. Mining changes across 103 real-world software ecosystems, WideSWE establishes 120 cross-repository tasks (balanced between 60 bug fixes and 60 feature developments) with comprehensive regression test suites. Across seven frontier agent configurations, full task success tops out at just 42.50%, exposing systemic failures in cross-repo dependency tracking and coordinated verification.

Key Takeaways

  • ✓Cross-Repository Benchmark Evolution: Moves beyond single-codebase evaluations by capturing interdependent multi-repository modifications across 103 real-world open-source software ecosystems.
  • ✓Frontier Agents Top Out at 42.50%: Across seven leading agent setups, end-to-end task completion ranges from 10.83% to 42.50% (Codex CLI + GPT-5.6-sol achieving the high mark), exposing acute bottlenecks in cross-repo dependency coordination.
  • ✓Isolated vs. Joint Workflow Dynamics: Proves that treating repositories independently leads to omitted cross-system dependencies, whereas joint execution surfaces essential multi-repo type signatures and test harness context.
  • ✓Complete Open-Source Test Harness: Benchmark code, environmental harnesses, and evaluation suites released at github.com/ZJU-ACES-ISE/WideSWE.
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点 / Background & Pain Points Autonomous software engineering evaluations like SWE-bench assess models strictly within single isolated repositories. In enterprise software ecosystems, however, feature implementations, API migrations, and dependency security fixes intrinsically span multiple interconnected codebases. Software agents must simultaneously coordinate changes across SDKs, services, and shared contract definitions. Existing benchmarks fail to evaluate this multi-repository operational competency. ### 架构亮点与底层机制 / Architectural Highlights Researchers from Zhejiang University introduced WideSWE, the first dedicated multi-repository coding benchmark: 1. 103 Software Ecosystems: Rigorously curated 120 real-world cross-repository tasks (balanced across 60 bug fixes and 60 feature developments) mined from interdependent public ecosystems; 2. Adapted Multi-Repo Test Harnesses: Restructured hidden test harnesses across multiple repositories to validate behavioral consistency and catch cross-package regression regressions; 3. Independent vs. Joint Execution Paradigms: Directly contrasts isolated per-repository workflows against unified joint-workspace execution to evaluate multi-repo context routing. ### 权威 Benchmark 与实测跑分对比 / Benchmark & Evaluation Benchmarking seven premier agent configurations revealed significant capability gaps: - Low Task Completion Ceiling: End-to-end task success spanned a narrow 10.83% to 42.50%, with Codex CLI paired with GPT-5.6-sol recording the top score (42.50%); - Prevalent Failure Modes: Traces highlighted that agents frequently fail to identify prerequisite secondary repositories, leave cross-package changes unfinished, or submit modifications that violate upstream contracts; - Joint Execution Superiority: Unified multi-workspace execution enabled agents to leverage downstream symbol definitions to guide implementation and execute bidirectional integration tests. ### 开发者实战落地与开箱指南 / Developer Practical Guide - GitHub Repository: Accessible at https://github.com/ZJU-ACES-ISE/WideSWE; - Paper Citation: Theoretical and empirical methodologies are detailed in arXiv preprint 2609.33382; - Tooling Architecture: IDE agent builders (Cursor, Cline, OpenHands) should prioritize multi-workspace root awareness and cross-project language server protocol (LSP) indexing.