Labeling code as AI-written increased visual attention, not measured defect detection

#topic

Labeling code as AI-written increased visual attention, not measured defect detection

Original: Khojah et al., Same Scrutiny, More Time: Eye Tracking Insights into Reviewing LLM-Labelled Code, revision August 30, 2026. Thirty-two software practitioners from 20 organizations each reviewed four shortened single-file Python changes in a counterbalanced, one-hour eye-tracking experiment. The selected code was all human-written, adapted from pre-2021 open-source projects; some segments bore a fictional AI-generation label, timestamp and prompt. This isolates perceived provenance, not the quality or review cost of actual agent-authored changes.

Labelled regions received longer gaze fixation: on a more complex 25-line segment, roughly 15 extra seconds, about 60% more than the corresponding unlabelled segment. Scan length—the authors' visual proxy for fine-grained inspection—showed no meaningful difference. Fourteen participants used the attached prompt in review, some as a requirement, others for comprehension. Caution: label and prompt were bundled, so the extra fixation cannot be attributed solely to the AI label; they deliberately did not score detection of injected faults or code smells. More gaze therefore cannot be called better acceptance, nor total active review minutes in a live workflow. The experiment tested neither merged PRs nor repairs.

Practical hypothesis: make request intent and the final patch available together, but keep a separate independently checked behavior oracle. Ask how fixation time, issue detection, corrections and review cycles change under a label-only versus prompt-plus-label test, rather than equating ‘more scrutiny’ with valid approval. See the clock comparison and acceptance artifacts.