Counterexample to ‘more structured team AI use always helps’: retail field experiment

#topic #adoption #counterpoint #field-experiment

Counterexample to ‘more structured team AI use always helps’: retail field experiment

Source: Alex Farach, Alexia Cambon, Lev Tankelevitch, Connie Hsueh, Rebecca Janssen, Scaffolding Human-AI Collaboration: A Field Experiment on Behavioral Protocols and Cognitive Reframing (Microsoft Research, April 2026; preprint linked from source). 388 employees at a Fortune 500 retailer had the same AI tool. A structured protocol requiring paired joint AI use was associated with lower document quality and substantially fewer documents than unstructured use; training that framed AI as a thought partner was associated with higher document quality toward the upper end of the distribution. Authors themselves flag AM/PM session confounding, differential attrition and LLM grading sensitivity to length, so do not present as a settled causal mechanism.

Relevance and limit. This is not coding-agent work, not a manager-led intervention and not a test of useful shared software-team artifacts. It is a warning about converting Dru’s useful-outputs/manager-participation idea into compulsory paired sessions or a use-frequency mandate: imposed collaboration can consume production time even when exposure increases. The proper coding-team test distinguishes free use with genuine outputs from compulsory joint tool use and checks independently assessed artifacts, downstream defects and human effort. Pair with adoption brief and registered team trial.