ParallelPilot: Supporting Coordination and Monitoring in Parallel AI Coding
ParallelPilot: Supporting Coordination and Monitoring in Parallel AI Coding
Long and colleagues first interviewed 14 frequent parallel-coding users about how they plan, isolate, log, observe and triage concurrent sessions. They turned those strategies into ParallelPilot: an editable dependency plan, Copilot session logs and a dashboard surfacing completion, errors, questions and stuck sessions. In a second, within-person test, 16 internal developers each got 20 minutes with the whole package and 20 minutes with ordinary Copilot on different six-ticket seed projects; the baseline could use plans, worktrees and parallel sessions. The package was paired with a short orientation outside the timed block.
Participants implemented 5.69 versus 4.19 tickets on average, and 14 of 16 versus 8 of 16 finished all six; peak concurrency rose 2.31→3.25 and self-reported tracking effort fell. Perceived control (3.94→4.31/7, adjusted q=.416) and perceived success at redirecting agents (3.94→4.13, q=.619) did not show statistically detectable improvements. The authors warn that progress status may distance developers from implementation details they need to judge and steer. A small null result is not evidence that dashboards cannot improve control.
These were twenty-minute tickets checked for implementation, not reviewed and merged PRs; participants did no systematic code review. No long-lived feature, active human review-time ledger, model-cost accounting, or post-release repairs. Formative observation was in June 2026 and the controlled trial in mid-August, despite the recent posting date. It suggests a testable interface question: does the supervision view let a human return from a status cue to the patch, original intent and integrated behavior before approving?