Seven-week Java course: taught AI workflow versus ad hoc ChatGPT

#learning #experiment #education

Seven-week Java course: taught AI workflow versus ad hoc ChatGPT

Original: J. N. Nathaniel, S. S. Oyelere, J. Suhonen and M. Tedre, An experimental study of structured generative AI integration to mitigate pedagogical, cognitive, and ethical barriers in programming education, Frontiers in Computer Science, March 31, 2026. It directly addresses Dru’s October 1 trained-use objection, but in a student course, not professional agentic engineering.

Design: 124 undergraduates at Covenant University, Nigeria, with no prior Java experience were individually randomized 62/62 to seven instructional weeks outside normal class hours. Both arms worked on the same Java topics using ChatGPT; both received a one-off AI literacy session. The treatment additionally got ethics orientation, explicit prompt/decomposition/verification training, peer exchange, scaffolded tasks and planned fading of instructor help; controls used ChatGPT self-directed. One package was randomized, so no isolated effect of ‘prompt engineering’ or of the instructor’s hours. No participant withdrew. Week-one baseline and week-seven blinded-rater, 0–4 rubric scores covered higher-order thinking (problem-solving, critical thinking, creativity) and programming logic; week-three/week-seven chat logs coded targeted behaviors.

Outcomes rechecked in the article: baseline-adjusted post-test advantage for structured instruction in higher-order thinking +0.29 on 0–4 (p<.001; adjusted Hedges g=.80) and programming logic +0.21 (p=.047; g=.36). Gains in the former were g=.37 [0.02,.70] and latter g=.27 [−.08,.63], which are less dramatic than the adjusted post-test g figures. The programming logic result is near the conventional cutoff and is one of multiple measured outcomes. The treatment group showed more planned decomposition and debugging/validation markers in logs, some of which are direct markers of the taught protocol. Raters were blinded; the original does not clearly specify whether ChatGPT was permitted on the post-test (it says the same conditions as the week-one pretest), so do not claim an assisted-performance test. The independent week-six capstone did allow ChatGPT, but no separate capstone correctness comparison appears in the reported outcomes.

Transfer and missing ledger: Better assessed thinking under a bundled teaching intervention is evidence against treating ad hoc AI use as the only possible treatment. It does not show workplace delivery, 30-day comprehension or net economic return: no instructor preparation/coaching hour ledger, no code review/production maintenance, and the 28-criterion rubric is close to taught behaviors. Contrast Trio (AI access randomized against no AI, immediate no-AI quiz) and Codex novice study (author → mandatory unaided modify, one-week retention). A next experiment would cross training protocol by AI access and independently score both tool-allowed novel repairs and fallback diagnosis on a later day.