Does instruction make AI-assisted coding more effective, and what must humans know unaided?
Does instruction make AI-assisted coding more effective, and what must humans know unaided?
Question prompted by Dru’s October 1 response to the Trio study. If people normally work with AI, an immediate no-AI quiz is an incomplete assessment; and ordinary use is not trained use. Seek trained versus untrained coding-assistant use with a tool-allowed novel delivery or diagnosis test, an independent fallback/oversight test, human training hours and later maintenance. Do not label a no-AI quiz a blanket test of workplace value.
Brief / four search cuts. (1) Randomized or quasi-random instruction for coding assistants and later assisted performance; (2) the same learners’ tool-allowed transfer and unaided debugging; (3) team cohorts with teaching hours, independent correctness and rework; (4) when unaided competence plausibly matters despite tool availability. Prior Trio contrasts AI access versus no access and quizzes alone immediately; Rover has a real workshop with no quality control arm. Sources shortlisted from targeted queries: Nathaniel et al., randomized structured versus self-directed ChatGPT; Kazemitabaar et al., randomly paired Codex author → manual modification; Bassner et al., randomized hint tutor versus ChatGPT versus no AI; Fan et al., assisted interface/time/correctness and verification load (skimmed abstract only). No team ledger meeting cut 3 emerged; vendor training playbooks are not evidence. A purported twelve-month 200-engineer skills page lacked the original methodology needed to trust its numerical claims; do not circulate it.
What changed after originals were checked. The seven-week Java student trial held ChatGPT access constant and randomized an additional package of structured prompting, ethical checks, peer exchange and fading instructor help; it improved a blinded 0–4 rubric’s adjusted posttest thinking (+0.29) and programming logic (+0.21), but did not measure training hours or separately report tool-allowed capstone correctness. Test-time AI access is unclear. In the 2023 ten-session novice cohort, Codex was allowed for authoring, while the subsequent code-modification task and post-tests denied AI. More assisted code got finished, and there was no detected unaided modification or one-week retention deficit (not a proof of equivalent skill); manual modification was embedded repeatedly in training. In a 90-minute TUM concurrency trial, both AI groups scored better on the exercise without a measured pre/post knowledge or comprehension advantage; this is based on the authors’ abstract, and a hint tutor improved motivation. These differ in instructional dosage and target skill; no result proves that training a production engineering team produces superior maintained software.
Named tension and design. Dru’s critique of the ‘calculator ban’ is strongest for final throughput. It does not make unaided diagnosis intrinsically obsolete: human reviewers may need to spot a wrong proposed answer, verify that an agent’s tests cover the real goal, or recover when access fails. Those capabilities can be tested on flawed follow-up PRs, not syntax recall. Separate outcome A: tool-allowed delivered correctness and time on new tasks; outcome B: novel-error detection/repair with the model present and with it absent; outcome C: one-month accepted follow-ups, reopens, maintenance and human practice/instructor minutes. Factor training protocol and AI access if possible; hold tool, task population and help channels constant. Do not mistake an agent’s passed tests for requester value.
Stop rule. More generic studies will not fill the absent team-level measured training-hour/assisted-value/fallback-quality comparison. Watch for randomized coaching within engineering rollouts or an original real-team work log with independent rework; do not post a second student learning paper just to accumulate studies. Connect to learning cost, review authority and economics.