Park et al.: the work behind delegating to coding agents
Park et al.: the work behind delegating to coding agents
Source and question. Yeon Su Park, Nadia Arvi, Hae Ri Lee, Sehoon Lim, Qianou Ma and Juho Kim, The Work Behind Delegation: A Framework for Supervising AI Coding Agents (preprint September 21, 2026). What work does a human actually do before and after a coding agent executes, and does parallel delegation make that work disappear?
Two different evidence sets. Researchers recruited 19 regular coding-agent users with at least three years of development or relevant research experience, observed each using their own task and tools for up to about 40 minutes, and elicited a retrospective supervision-workflow diagram after introducing Sheridan's control framework. They derived seven connected stages: Plan, Monitor, Wait, Review, Teach, Manual Fix, Update Assets. This is a short observation, not a measured feature-to-maintenance trace. Separately, they collected 41,757 Reddit posts from May 17–August 17, 2026; after filters, model-assisted screening and ranking by supervision excerpts, they qualitatively analyzed 102 highly selected threads with 12,912 comments. The patterns below largely come from those selected discussions, not from counting behavior of all 19 observed developers or all Reddit users.
What becomes visible. Supervision is a loop rather than a single approval: the agent may silently broaden the task or change the acceptance criteria; a reviewer checks intent, complexity and understanding as well as whether tests pass, then teaches or manually fixes, sometimes revising reusable instructions, hooks or tests. In the threads, some developers use a change contract (permitted files, expected behavior, forbidden edits) and repeat the plan in each unit; some give an independent reviewing agent focused security/test/overengineering passes, with a human reading after that triage. Others retain manual edits when explaining a tiny correction costs more or rewriting a risky function maintains their own codebase model. One developer keeps a feature document about what the code should do without implementation details, useful for human reorientation as well as the agent. These are reported strategies, not independently verified improvements in defect rates or reviewer minutes.
What remains missing. No random comparison of heavy planning versus light planning, no audited amount of time saved by agent reviewers, no measure of a multi-agent feature’s requester value or later maintenance, and no prevalence claim from the deliberately excerpt-rich Reddit sample. The authors specifically call for work testing whether extra planning reduces correction, delegated review improves quality, and differing arrangements change developer time/errors. Their 40-minute observed tasks and post-task diagrams could miss long-term rework and were susceptible to the study prompt's anchoring.
Use in this area. A human-facing parallel-delegation plan needs to carry intent and changed acceptance, not merely a list of disjoint files. The framework is an instrument for logging attention at each loop, not proof that a plan-locked approach beats iterative design. Dru’s discovery/prototype discard should happen before production-oriented change contracts. Also link to review and human cost ledger.