WideSWE: a repository can pass while the feature still fails
WideSWE: a repository can pass while the feature still fails
Source/date: Baoyi Wang et al., WideSWE, posted September 27, 2026. New experiment on historical multi-repository work: GitHub PR corpus of merges from January 2024 onward; eligible case’s latest PR merged June 2025 or later. This is not a live September deployment. Question: When one request requires two or three repositories, does an agent keep the full feature and cross-repo interface intact?
Selection: authors gathered 1,729,171 PRs from 103 GitHub ecosystems, filtered linked cross-repo candidates, manually reviewed 635, validated 192 and selected 120 (60 fixes, 60 features) from 41 ecosystems. Prompts combine original issues and PR descriptions. Hidden tests from historical patches were reviewed against request intent and adapted when they demanded implementation details or extra features not actually requested; the paper does not give a count of revised cases in its main text. Case-level success requires all fail-to-pass and regression tests in all target repositories, not a per-repository average. Seven model/scaffold configurations solve 10.83–42.50% of full requests; the best Codex CLI/GPT-5.6-sol solves at least one target repository in 83.33% but all in only 42.50%. In 48 of its 69 unsuccessful cases, at least one repository passes its new-behavior checks and another does not.
Joint versus isolated under identical prompts: 89 cases, Codex CLI/GPT-5.6-sol: one joint ecosystem run passes 36/89 (40.45%); independent runs per target pass all targets in 32/89 (35.96%). There are 20 outcome reversals: isolation rescues missing work in some fixes but loses cross-repo behavioral context on many features. Independent agents consume 296.6 mean API requests per task against 94.7 joint—each independent run gets a full per-run budget, not matched total task spend. On 60 features, success falls 43.33% joint→31.67% independent; on 29 fixes, the direction reverses, 34.48%→44.83%. These are descriptive task-sample effects, not an unconditional rule against parallel agents. Sentry Go/Python/Ruby shows joint work narrowing to Go while independent runs cover all three; Ansible and Symfony show benefits of shared cross-repo context.
Boundaries: generated tests do not establish live requester value, semantic completeness beyond tests, review cost, merge, maintenance or human time. Independent per-repository runs are not agents synchronized by a human planner or a shared-contract document. Contrasting this with ParallelPilot's small-ticket coordination gain points to a hypothesis about whole-feature coverage, not an apples-to-apples refutation. Link to recent feature brief, planning and feature outcome. Hold-back: post when whole-feature success across service/repository boundaries is the focus, or when the ParallelPilot post has landed; do not stack adjacent September 27 papers now.