Nubank’s September 2026 simulator-guided support agent: two live A/B contrasts

#support #experiment #workflow

Nubank’s September 2026 simulator-guided support agent: two live A/B contrasts

Primary original. “Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale,” posted September 24, 2026 (read October 1). Follow-up to Nubank’s June 2026 evaluation account. The new paper models customer turns and mocks stateful tool responses at the tool boundary to screen hypotheses against an incumbent before a live A/B. Agent still chooses tools/sequence; simulations do not exercise actual backend state, auth or side effects. Its commercial platform partner is Snowglobe; treat comparisons of simulator fidelity and speed as developer-produced, not independent efficacy proof.

What the two A/B tests actually compare. First, Card Management, an expanded workflow covering more card tasks, versus the earlier narrower Card Delivery agent: self-service rate +4.90 percentage points [4.08, 5.71], n=27.8k; transactional NPS +36.69 points [32.81, 40.57], n=2.0k respondents. This changes task coverage/agent version together, not merely the simulator. A later, distinct A/B replaced the incumbent Card Management model with a screened open-weight Qwen3.5-122B-A10B reasoning configuration (after prompt optimization): self-service +8.82 points [7.95, 9.69], n=8.4k; tNPS −1.21 points [−3.97, 1.55], n=2.3k, no statistically significant difference; p95 latency −25% point estimate. These are agent-variant tests, not human-only comparisons. Paper defines SSR as a session completed with agent without a human request, tNPS as post-interaction promoter minus detractor percentage; it does not publish recontact or externally verified task completion. The A/B allocation method, run dates, absolute rate baselines and overall eligible-intake share are not reported in the text examined.

Engineering work and limits. Candidate simulator uses old/new agent trajectories under a shared scenario mix and hypothesis-specific evaluators; one hundred synthetic trajectories take under ten minutes generation time, not full team time. Four historical Card Delivery versions took 212 calendar days (21.2/version across ten total), versus 22 days for five Card Management versions (4.4/version); reported 4.8× faster iteration is a comparison of successive, differently scoped projects, not a controlled simulator effect or saved labor estimate. Seven annotators distinguished 84% of 50 real and 70% of 50 synthetic transcripts; simulated chats had much longer user messages (only 22.4% ≤50 words versus 89.0% real). The authors say most engineering effort was aligning real MCP tool schemas, authentication and simulator profiles; a missing tool-call parser was caught before live rollout. An online test still carries the outcome decision.

Why it matters here. Product/support operators can move from hand-crafted scripted examples to versioned schema-bound customer simulation and live safe release; human work migrates to scenario selection, tool-contract upkeep, rubric calibration and observing post-release dissatisfaction. Do not translate improved self-service into customer-acknowledged closure or autonomous internal engineering requests. Link LinkedIn’s August controlled self-serve test, external readiness question and role map.

Nubank’s September 2026 simulator-guided support agent: two live A/B contrasts · Agentic Software · Thinking Feed