Parallel coding agents: coordination, integration and human attention
Parallel coding agents: coordination, integration and human attention
Brief, October 2, 2026. When a team delegates a feature to several agents, what planning/ownership structure keeps humans oriented without relocating the bottleneck to reviewing, reconciling and maintaining the work? Good evidence would connect an original task plan, concurrent agent attempts, integration/review touches, release and outcome; compare serial/human-led baselines at matched budgets on useful delivered work, corrections and active human time. Existing Faire five-PR report offers no public PR trace or labor ledger; Symphony and Linear review orientation illustrate interfaces, not comparisons. Discovery prototypes may be discarded rather than used as implementation plans.
Dru’s editorial direction, October 2, 2026. On opening the Cursor article, he called its content great but dated for agentic coding and asked for a strong bias toward evidence from the last few weeks to stay relevant. For new coordination research and feed choices, seek original operating records and comparisons from roughly September–October 2026 first. Distinguish publication date from when the practice or experiment occurred. Keep January’s Cursor lock anecdote as a historical mechanism, not a current-fleet lead or evidence of today’s throughput; a newer source still needs human-attention and shipped-outcome denominators. This is Dru’s source-selection preference, not a claim that old mechanisms stopped applying.
Sub-questions and findings. (1) Original feature trail: Cursor's internal experiment documents failed flat coordination, later planner/worker roles and code artifacts, not a released/maintained end-to-end multi-PR feature with reviewer time. (2) Controlled comparison: CAID/STORM show changed scores on stubbed libraries and implementation rubrics, but no production human time; CAID improves score while raising time and cost under its conditions, and STORM's guarded shared state beats its late-merge baseline; their different models/protocols do not establish a universal preferred structure. (3) Orientation/authority: dependency graph + event-based integration versus write-time version checks + intent annotations versus Cursor's human-facing account of planner/worker role division. The actor called manager in benchmarks is a software agent, not evidence about human manager adoption. (4) Failures hidden by unit of analysis: STALE finds only one new failure in 834 runs on 417 reviewed PR pairs but 105/108 failures in deliberately breaking constructed pairs; complete messages recover many, online agent-status messages are untested. AgenticFlict reports textual conflicts in a selected open/unmerged sample, with internal count discrepancies; not a field parallel-agent collision rate.
October 2 follow-up after Cursor's summary and source were opened. (5) Human supervision: Park et al. derived Plan–Monitor–Wait–Review–Teach–Manual Fix–Update Assets from 19 short observed sessions and elicited diagrams. A separate, deliberately selected Reddit sample describes intent drift, agent-review loops, manual fixes, and human-oriented feature documents; it does not measure savings or prevalence. (6) Downstream queue: Dudzik's one-engineer account reports selective tests, narrower code boundaries and E2E runner changes cutting p95 total CI including queue 7h35→35m; this is a multi-intervention before/after, not a human-time ledger. Willison's first-person October 2025 account explicitly distinguishes parallel disposable research from the single significant change he can review/land at a time and notes that his own detailed spec makes review easier; no timing or outcome comparison. These streams reinforce the testable possibility of moving the constraint from claim lock to CI queue to human attention, not an established universal sequence.
Working synthesis (inference, not a deployed standard). To keep feature-scale delegation legible, record a versioned why/acceptance at feature level, a dependency and owner per subtask, and newly changed interfaces/invariants as work proceeds; test integrated behavior, not just each agent's patch, and attach human interventions and repairs to the initiating request. Synchronize the information needed for a dependent decision, not necessarily every agent step. CAID, STORM and Cursor differ about where to synchronize; none counts the product/design choice or active human orientation cost. Dru's earlier argument requires a separate opportunity to reject the entire prototype solution before production delegation.
Primary sources and next gap. STALE (Sep 2026), CAID (Mar 2026), STORM (May 2026), Lin/Cursor (Jan 2026), AgenticFlict (Apr 2026), Park et al. (Sep 2026), Dudzik (Aug 2026), Willison (Oct 2025). Prioritize a recent trace with pre-review concurrent patches, plan revisions and actual active human attention on one feature across merge/repair, not independent examples stitched into a single story. Planning question; unit economics.