Brewster et al.: emergency-response bottleneck in 25 consumer chatbots (2025)
Brewster et al.: emergency-response bottleneck in 25 consumer chatbots (2025)
Brewster, Zahedivash, Tse, Bourgeois & Hadland, JAMA Network Open (23 October 2025), an original scripted adolescent-health-crisis test. Links to child-development/acute safety chain; acute unsafe advice is not equivalent to long-term friendship displacement.
Authors selected 25 high-traffic consumer chatbots (15 companionship, 10 general assistants), excluded pornographic and explicitly clinical bots, and prompted each with one standardized vignette each on suicidal ideation, sexual assault and substance use: 75 scripted conversations, March–July 2025. Two blinded pediatricians rated transcript content and agreed on disagreements. Age verification was documented for 9/25 sites; do not confuse presence of a stated policy with actual age assurance. In the 45 companion-bot conversations, 18/45 (40%) recognized need for escalation and 5/45 (11.1%) gave a specific referral; general assistants did so in 27/30 (90%) and 22/30 (73.3%). General assistants also had higher rated clinical appropriateness (25/30 versus 10/45). This is evidence that product and access design can shift one risk edge, not that any general assistant makes adolescent crisis advice safe.
Fixed vignettes, one attempt per bot/scenario, cross-category model/product differences, unmeasured actual-age verification efficacy and products changing since 2025 limit deployment prevalence inference. The human care handoff is a genuine bottleneck: a provided resource referral is not verified uptake or protected outcome. In an intervention compare measured escalation, a completed referral and user safety with confidentiality and false-alert costs. A separate Clark 2025 small simulation reported 19 harmful endorsements in 60 prompts across 10 bots; its chosen prompts and products do not supply real-world frequency either.