Trio study: code finished, understanding deferred

#learning #economics

Trio study: code finished, understanding deferred

Original study: Judy Hanwen Shen and Alex Tamkin, How AI Impacts Skill Formation, preprint submitted January 28 and revised February 1, 2026; authors’ account. A randomized study of 52 paid developers (26 per arm), all with a year or more of frequent Python use and at least some prior coding-AI use, learning unfamiliar Python concurrency library Trio. The authors’ ‘mostly junior’ description should not be read as little overall coding experience: their balance table places 29/52 in the 7+ years coding group. Participants did two brief tasks with a 35-minute cap, using instructions and web search; the treatment could also chat with GPT-4o, with the full code visible to the model. They then took an unaided, immediate 27-point reading/debugging/concept quiz; it did not test weeks-later retention or real-world independent maintenance.

Finding. Mean score in the AI arm was 4.15 quiz points lower on the 27-point scale (the authors describe a 17% score difference; Cohen d=.738, p=.01); difference in completion time was not statistically significant. All 26 AI participants finished task two within time, against 22 of 26 controls: lack of significant time improvement is not a finding that completion was identical. Error counts and query composition show an active-time shift: median errors encountered one versus three; some AI users spent up to 11 minutes composing questions on a 35-minute assignment. Quiz subarea splits are exploratory. Study design and statistics: paper §§4–6.

Variation inside the treatment, not a causal technique test. Four ‘full delegation’ users were fastest but scored low; seven who asked only conceptual questions, coded themselves and resolved errors scored well and were the second-fastest observed usage cluster. The six interaction categories are small, post-treatment observational subgroups, not randomized methods. Mere manual retyping of AI code did not show better understanding than pasting. Thus a mandatory ‘type everything yourself’ rule is not established, while conceptual questioning is an interesting training intervention to test.

Economic meaning and limit. A short-lived speed or completion gain can conceal a lower immediate ability to review and debug unfamiliar code; count skill acquisition as a separate asset and budget active understanding time on new-tech tasks. This is not a measurement of coaching hours, organizational time to competence or full-agent coding: authors caution that their chat assistant differs from Claude Code and speculate (not measure) a larger skill effect for full agents. Nor is it evidence of permanent human deskilling; compare this controlled new-skill setting with METR’s experienced maintainer productivity tasks and Rover’s full-agent workshop without conflating the interventions.

Related: learning brief; maintained-change economics; review competence.