Self-evolving Agentic Customer Support System at LinkedIn

#summary

Self-evolving Agentic Customer Support System at LinkedIn

The original LinkedIn system periodically generates candidate prompts under fixed business-policy rules; evaluation agents check grounding, intent, and multilingual quality against human-labeled checks, then gate and stage changes as versioned artifacts. It uses agent-invoked knowledge retrieval and can roll back or fall back to human support. This is a change to support and engineering roles: maintain authoritative articles, tool policy, evaluator rubrics and release gates rather than just hand-writing every answer or prompt. The intervention bundled these changes, so the A/B cannot assign credit to prompt evolution versus retrieval or evaluation separately.

The live experiment assigned users once 50:50 for two weeks, comparing an existing agent with the integrated workflow. QA and cancellation arms contain about 89,000 and 41,000 conversations respectively. The reported self-serve metric concerns avoiding human escalation or completing a cancellation without handoff; the paper says satisfaction and moderation did not regress but does not give the rates, customer-confirmed outcome, repeat contact, wrong-action counts or human effort. Routing accuracy is from 356 labeled decisions per condition, even though it is presented beside the live results. This contrasts with Taobao’s older human-versus-AI-eligible test, which has a different baseline and task mix; neither comparison alone tells us whether a given customer’s problem stayed solved.

Read at arxiv.org · 18 min