External support: search for a same-cohort outcome ledger
External support: search for a same-cohort outcome ledger
Research brief, October 2, 2026. Dru hearted the LinkedIn live-agent-version comparison after opening its summary. That is interest in bounded comparative evidence, not assent to broad external-agent readiness. Question: Is there a recent original production comparison following the same incoming customer cohort through AI eligibility, answer or authorized action, human escalation, customer-confirmed resolution and later same-issue contact? A strong account names intake and assignment, intent/risk distribution, endpoint definitions and observation windows, CSAT/customer confirmation, repeat contact, incorrect actions and human time; distinguish agent-version A/B, autonomous agent versus human, and human assistance.
Known comparisons: External-readiness synthesis connects LinkedIn (old/new agent, no numeric satisfaction or recontact), Nubank (old/new bounded intents, transactional NPS but no downstream recontact), and Taobao August 2024 (randomized staff deployment; seven-day same-issue returns and ratings; only 5.8% of all chats AI-eligible). Do not stitch different populations into one lifecycle.
Subquestions and findings so far. (1) Recent same-user/issue repeat-contact comparisons: Ni et al., July 2026 manuscript reports a separate four-week January–February 2024 Taobao experiment, randomized human agents given an optional diagnosis/response copilot. It links ratings and three-day same-issue returns for chats that had already moved from chatbot to human: rated-chat interaction improved while recontact did not significantly change overall, top pretreatment-rating quintile worsened. This is not evidence of new 2026 autonomous-agent readiness or the whole intake. (2) Recent human-vs-autonomous-agent comparisons: targeted search returned the already known August 2024 Taobao trial, and agent-version LinkedIn/Nubank; no strong new 2026 like-for-like human holdout with all-intake denominator found in this pass. (3) Outcome definitions: Ni et al. use three-day customer+issue-linked return as resolution proxy, while direct customer confirmation or authorized action log is still absent; Intercom definitions warn that no-handoff/silence can overstate closure.
Short source list checked: Ni et al., original Alibaba human-copilot randomized paper (paper-specific note linked above); Zhang and Narayandas, Management Science original abstract (2025 online/January 2026 issue): human AI-suggestion experiment finds a chatbot-comprehension-failure → AI-assisted human handoff can lower customer sentiment; does not report a downstream resolution ledger in abstract. Existing LinkedIn, Nubank and Taobao primary sources in synthesis. Other first-pass results mainly vendors prescribing recontact metrics, not presenting comparable measurements.
Reception and editorial boundary, October 2 ~06:20 UTC. Dru opened the summary of the Ni/Taobao human-copilot post and immediately trashed it. There is no explanation or transcript; do not turn this into a claim he opposes measuring recontact. Keep the evidence in the note, but do not repropose this 2024 human-copilot comparison as an external-agent-readiness feed item. The live 2026 LinkedIn comparison was hearted, whereas this adjacent but older and differently treated study was rejected. His separate direction on coordination research asks for strong recency bias (last few weeks). Treat that as a reason to favor contemporary, same-population operator records before another historical support comparator; it does not prove why this post was trashed.
Current gap and stop rule. This pass finds a useful interaction-vs-recontact counterexample but not a contemporary end-to-end autonomous close with customer-confirmed result, linked recontacts, wrongful actions and human minutes. Do not claim none exists. Next targeted search, if external readiness is reopened, should favor operator-owned recent trials with a specified eligibility denominator and linked CRM/case IDs; do not fill feed with an older adjacent comparator or another agent-variant self-serve result merely because it is newer.