CAID and STORM: when to synchronize parallel agent edits

CAID and STORM: when to synchronize parallel agent edits

Primary sources: Geng and Neubig, Effective Strategies for Asynchronous Software Engineering Agents (CAID), March 2026; Liu et al., Multi-agent Collaboration with State Management (STORM), May 2026. Both use Commit0-Lite (16 stubbed Python libraries) and PaperBench Code-Dev (implementation rubric, not full reproduction); neither measures shipped products, actual human review/correction, post-merge maintenance, or user value.

CAID mechanism and comparison. A manager makes a task dependency graph, assigns work to isolated git worktrees, merges finished commits one at a time, tests and reassigns; the originating agent resolves merge conflicts. For Claude Sonnet 4.5 on Commit0-Lite, mean score 53.1 single → 59.1 with four agent engineers; reported runtime 692.6→1583.2 seconds and inference cost 1.9→8.1 in table units. For PaperBench Code-Dev, score 57.2→63.3 with two engineers; runtime 1803.5→2080.4 seconds and cost 3.3→6.5. Extra workers raised correctness, not faster wall-clock in these conditions. On Commit0-Lite, eight workers did worse than four; integration and task overlap grew. No human effort counted. Caveat: different per-configuration iteration budgets, so this tests bundled organization plus greater compute, not parallelism isolated at fixed spend.

STORM mechanism and comparison. Manager also allocates initial scopes, but workers share one workspace; file version checks reject writes when a file the writer read changed, returning diff and stale dependencies; in-code annotations signal intent. With Sonnet 4.6 on Commit0-Lite, macro test pass 82.5 STORM vs 63.8 isolated GitWorktree vs 66.4 single; total-test weighted pass 46.2 vs 24.6 vs 20.7 respectively. Authors say single agent wins cost-efficiency across models; raw cost rises about $199→$429 as STORM's maximum agents rise 2→8 across 16 repositories, while wall time stays roughly flat. Under their STORM protocol, 8 max improved outcomes relative to 4, especially on two high-test-volume repositories. Its worktree comparator used the earlier architecture as a baseline, but changed model (Sonnet 4.6 instead of 4.5), implementation and protocol; do not compare scores across the two papers as a replication.

Named disagreement / resolution limit. CAID concludes dependency-aware isolated worktrees and explicit merge gates stabilize parallel implementation; STORM reports that its own write-time shared-state checks outperform a late-merge worktree baseline and reverse CAID’s four-to-eight-worker pattern under its task distribution. Neither establishes a universal agent-count optimum or that naive shared folders are safer than isolation: STORM’s shared state is guarded by read-dependency versioning. STALE independently shows why textually clean merge is not a behavioral guarantee; file-version checks also need cross-file invariants or combined tests. Next experiment should match model, token/cost budget, task mix and integration/tests, randomize coordination, and record active human orientation and repair time. Coordination brief; economics.