ParallelPilot: faster ticket coordination did not establish better steering
ParallelPilot: faster ticket coordination did not establish better steering
Primary source: Tao Long et al., ParallelPilot: Supporting Coordination and Monitoring in Parallel AI Coding, posted September 27, 2026. Observation dates: 14-person formative interview study mid-June 2026; 16-person controlled probe mid-August 2026. Recent paper, several-weeks-old trial. It tests a supervision interface, not fleet-scale maintained-change throughput.
Design: Interviews and screen-shared workflows surfaced PILOT: Plan, Isolate, Log, Observe, Triage. The probe combines dependency-aware planning, a Copilot run logger, and an ambient dashboard for agent state and intervention cues. Sixteen internal developers/researchers, experienced with assistants but mostly self-described novices in parallel coding, did two counterbalanced 20-minute coding blocks on different seeded projects. Each had six mock tickets with a shared-file pair, an ordered dependency pair, two independent tickets, and a product/policy choice. Baseline Copilot CLI/Chat users could plan, use worktrees and open multiple sessions. The probe arm additionally received a five-minute onboarding video plus trial setup outside the timed coding window. The whole interface-plus-orientation package differs, not one isolated feature.
Findings: With lightweight recording/code checks of implemented tickets, probe participants completed mean 5.69/6 versus 4.19/6 baseline; ticket throughput was 0.445 vs 0.272/min (+63%, BH-adjusted q=.007); peak concurrent sessions 3.25 vs 2.31; 14/16 finished six tickets vs 8/16. Mean elapsed time to finish tracked tickets fell 16.33→13.76 minutes across the study. In the eight participants who completed all six in both blocks, plan/isolate accounted for 25% of task duration with the probe versus 27% without; both phases shortened. On seven-point self-ratings, tracking effort dropped 4.63→2.75 and status-cue awareness rose 3.63→6.19. Yet perceived control 3.94→4.31 (q=.416) and reported success redirecting agents 3.94→4.13 (q=.619) had no detectable difference; nine participants rated verification ease equal. One baseline-preferring user said summary abstractions reduced contact with implementation. This is a null detection in a small study, not proof the interface cannot help steering.
What remains unmeasured: no PR/review/merge required, full acceptance/behavioral verification, subsequent production/repair, or requester result. The probe reports elapsed time and phase shares but not attributable active human attention per integrated feature, model spend or full ownership cost. Six-ticket ceiling, twenty-minute blocks, controlled model/configuration and internal recruitment constrain external validity. Six of eight who saw the probe first said they reused its planning lessons in the later baseline: order counterbalancing does not fully exclude carryover. An actionable next test gives humans a direct path from status back to the reason, patch and integrated behavior needed to decide whether an agent’s approach is right. Connect coordination brief, planning, Park, feature-to-outcome and whole-feature cross-repo benchmark.