Nubank’s bounded support agents: A/B gains against older agents
Nubank’s bounded support agents: A/B gains against older agents
Original study: Aman Gupta et al., Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework, KDD 2026 industry paper by Nubank practitioners. Five production use cases with online A/B comparisons against earlier agent/legacy support variants, not randomized all-AI versus all-human support. Card delivery: +37 percentage points in AI transactional NPS, +29 points self-service rate relative to earlier variant; tNPS remains 10 points below expert human agents (nonrandomized comparison). Across other four cases tNPS gains +4.5 to +40 points; self-service changes −1.5 to +7.5 points; debt management still 23.6 points below human-agent tNPS. Authors do not show active human time, true all-in labor savings, public customer traces or longer-term same-issue recontact by arm in the paper.
Mechanism actually changed: from a knowledge-base retrieval and transfer-only initial variant, the card-delivery agent acquired standardized routines, logistics/customer-data tools, frustration detection and proactive transfer, and self-contained card-reissue actions. Agent versions were gated with human-labeled offline evaluations, staged at 1% of traffic then expanded. The authors favor narrow tools, idempotency and human handoff below confidence bounds; they say no unconstrained or irreversible credit or debt decision is given to agents. With better tools and evals, both self-service and user satisfaction can improve together within eligible work; it does not settle manager attention or software-change ROI.
Relation to Dru’s September 30, 2026 thesis: evidence for moving beyond agent summaries to bounded action, with tested rubrics and safer escalation, but not proof of autonomous resolution across the whole incoming queue. Compare Taobao’s randomized worker trial (different place, generation, population and comparator), Intercom’s resolution definitions, and brief.