Four to five days of agent use: a threshold in PR association, not yet a threshold effect
Four to five days of agent use: a threshold in PR association, not yet a threshold effect
Source and measurement. Murphy-Hill, Butler and Savelieva, Adoption and Impact of Command-Line AI Coding Agents, §5.2.1, Fig. 7 (July 1, 2026). Within an early-2026 panel of active Microsoft agent adopters, compare an engineer’s merged-PR counts in weeks with 0, 1, 2, 3, 4 or 5+ tool-use days, with engineer and calendar-week fixed effects. Relative to the same engineer’s zero-use weeks, the point estimates are approximately +3%, +5%, +15%, +22%, and +50% PR counts. The top bin is five or more days in a calendar week, not a measure of daily use over a month. The 4-to-5+ gap (+22% versus +50%) is an association, not an identified causal discontinuity: tool use may be endogenous to deadline, task load, extra workdays, task selection, or changes in workflow; fixed effects do not remove those varying circumstances. Five use days need not mean weekend work, but the 5+ bin may include six-/seven-day weeks. PRs are not customer value or net productivity.
Dru’s interpretation, September 29, 2026. The sharp rise may reflect crossing from light assistance to a new working practice. Dru connected this to a Netflix training report emphasizing an ‘aha’ moment rather than occasional use and to Tessl’s reported lesson: commit to teaching agents to complete work correctly rather than reverting to personally fixing their mistakes. These comparisons are Dru’s interpretation, not independently measured explanations for Microsoft’s intensity pattern. On a Tessl-hosted panel, Tessl’s Patrick Debois advocates fixing the system that produces agent-generated code rather than repeatedly fixing its outputs; this aligns with the proposed mechanism but is not a quantitative test of a 5-day threshold. The available Netflix rollout abstract does not itself report the training comparison. In the Microsoft PR-output post discussion that ended September 29, Dru and the group explicitly resisted treating use frequency as a target: workload or weekend work could also account for some of the high-use-week association. A manager-led adoption strategy should measure useful team outputs, not force five-day use.
Research cut. Disaggregate exactly five weekday-use days from six or seven, logged work hours, pre-week task load and PR size/complexity, alongside quality, active review time, rework and inference spend. To test the learning account rather than simple volume, look for later output gains at similar usage intensity after a documented shift to agent-first delegation and better reusable team constraints. Keep separate the question ‘how do engineers reach sustained use?’ from ‘does sustained use improve net value?’
Related: team-first adoption; productivity evidence; unit economics.