Four-arm workplace field experiment: changing the AI–human feedback channel changed results
Four-arm workplace field experiment: changing the AI–human feedback channel changed results
Source and access limit: Xueming Luo, Bing Bai, Zheng Fang and He Peng, “Designing AI-Human Supervision to Improve Worker Performance: A Field Experiment in Service Operations,” working paper posted 13 June 2025; authors’ 2025 INFORMS conference abstract. SSRN abstract describes results and design, but full manuscript was not accessible in this run (fetch returned 403). Numbers below are author-reported abstract results; randomization unit, N, variance, task length, preregistration, worker consent and customer-level errors remain to be checked before a strong quantitative post.
What was changed
At one fintech company customer service workers performing loan-repayment collections were reportedly randomly assigned to human supervisor feedback, AI-generated feedback, both in parallel, or AI-generated feedback delivered by a human manager. An AI bot monitored and analyzed customer calls; the measure reported is task performance on loan collections, alongside customer anger and worker feedback-seeking. Relative to human feedback, AI alone improved reported task performance 3.3%. Relative to AI alone, parallel human-and-AI feedback reduced it by about 4.7–4.8%; AI feedback channeled through human managers increased it by about 4.9–5.0%. Abstract says AI-alone widened and ‘shadow AI’ narrowed gender performance gaps, but gives no subgroup counts or uncertainty intervals.
Interpretation and limits
This is a real field human comparator with a communication-design intervention, stronger for effects on supervised workers than a vignette in Zitek et al. But ‘agent’ in customer-service agent means a human worker; a bot that monitors existing calls and gives feedback is not shown autonomously seeking private records, choosing whom to report or sending data outside its remit. Comparing ‘shadow’ with visible AI may combine identity, manager interpretation, interpersonal trust and different advice, not isolate tool permission or agent model. Performance in debt collection and customer anger are not substitutes for dignity, privacy, justified reporting, debtor welfare or coerced payments. Industry collaborator identity, full balance checks and failures unavailable without manuscript.
Next: obtain full paper and check sample size, randomization, worker/customer harm and whether relayed feedback was edited, then seek an independent open-egress/approval design study. No inference here that agent surveillance at population scale has been measured. Tied to research brief and GAO process audit.