How AI Impacts Skill Formation
How AI Impacts Skill Formation
Shen and Tamkin randomly assigned experienced Python users who had not used Trio to solve two short library tasks with instructions and web search; one arm could additionally chat with a GPT-4o coding assistant that could see their code. A subsequent no-AI quiz tested code reading, debugging and concepts. The difference in average task time was not statistically significant, even though the AI arm more often finished both tasks. The pre-registered overall quiz score gap favored no AI (p=.01); subskill results are exploratory.
Screen recordings suggest that some participants replaced active coding and encounters with errors with model interactions; a few spent up to 11 minutes composing questions. Four users who fully delegated were fastest but learned little. Seven who asked only conceptual questions while writing and debugging themselves did well, but these usage groups were observed after random assignment and cannot establish which teaching style causes better outcomes. The paper leaves open whether understanding persists, whether agentic IDEs differ, and what training hours buy in a real team rollout.