Taobao human-support copilot: warmer chats without fewer same-issue returns

#topic #support #experiment #roles

Taobao human-support copilot: warmer chats without fewer same-issue returns

Original source. Xiao Ni, Yiwei Wang, Tianjun Feng, Lauren Xiaoyuan Lu, Yitong Wang and Congyi Zhou, Generative AI in Action: Field Experimental Evidence from Alibaba’s Customer Service Operations, July 2026 manuscript. Alibaba randomly assigned 5,940 human after-sales agents with tenure under one year (2,895 treated, 3,045 control) to four weeks of access, January 23–February 19, 2024, after four untreated weeks. Some 2.56 million chats took place during treatment. Unlike Taobao’s later August 2024 autonomous-eligible experiment, this is a human-copilot intervention: the customer had already passed through a chatbot and sought human support; the assistant suggested an issue diagnosis and solution at the beginning, which the human could send, revise or reject.

Same incoming cohort, two imperfect outcomes. The randomized access effect, estimated in an agent-day panel, reduced issue-identification time approximately 8% and duration approximately 1%; mean rated-chat score rose 0.042 on a five-star scale, dissatisfaction (one or two stars) fell 1.2 percentage points. Three-day same-customer, same-issue retrials did not significantly change (point estimate 0.000). Approximately 38% of chats generated a same-issue return within three days in the descriptive sample; only about 15% received ratings, though response fractions appeared stable between arms. These are linked measurements of customer experience and behavior, not a customer-confirmed resolution measure; a return can reflect many frictions and no return can conceal an unresolved issue. Suggestions were sent or edited in only 21.5% of treated chats; the usage-IV estimates apply to compliers and require stronger assumptions than the assignment comparison.

Human attention is uneven. By pretreatment rating quintile, bottom-quintile access raised mean ratings by 0.812 and reduced dissatisfaction by 21.5 percentage points; top-quintile access lowered ratings by 0.283, raised dissatisfaction by 7.0 points and raised three-day same-issue retrials by 0.9 point (p<.05). Do not mistake this Q5 result for the average or for every veteran; the sample excludes workers with over a year’s tenure. The authors propose a workflow-disruption explanation: treated high performers switching across concurrent chats took longer to return and reply later in a conversation; immediate retrials increased. That mechanism is suggestive observational process evidence, not randomized isolation of attention switching. The study’s matching algorithm used historical performance for case assignment, so between-quintile task mixes need not match.

Read against LinkedIn’s external-agent readiness. A properly randomized improvement in rated interaction need not reduce a same-issue return; LinkedIn’s later old-versus-new agent measured self-service but omitted both numeric customer satisfaction and recontact. Neither paper is a universal readiness answer. This human-assistance trial shows why a support-builder should inspect downstream outcomes and how the person’s attention is displaced, not simply a fast first response. Its experiment was in 2024, so the July 2026 manuscript is not a contemporaneous 2026 deployment test or evidence that newer autonomous agents fail. Also see outcome-ledger search and cross-functional roles.