Child/teen AI companions: what outcome evidence tests developmental harm?

#topic #child-development

Child/teen AI companions: what outcome evidence tests developmental harm?

Research brief, 2 October 2026. Question: does repeated companion chat with AI change children’s social development or safety relative to feasible alternatives? Good evidence would track who actually uses which features, whether they seek unavailable human help, the short- versus long-run validated outcome and whether age assurance, bounded relationship language and crisis handoff work in use. Gates’s August essay calls evidence small and mixed. This is a product-deployment/social pathway distinct from rogue AI and from state biometric search.

Actor, edges and original tests

Provider chooses relational persona, memory, rewards for engagement and age/crisis policies → minors gain access (age labels do not enforce themselves) → perceived support or deference → disclosure, peer displacement or crisis advice → longer-run distress and development. Common Sense/NORC’s 2025 survey measures the access/use edge: 72% ever tried a broadly defined companion, 52% a few times per month or more, but among users 80% still spend more time with friends and only 6% more with companions. It cannot attribute social change to use. Kim et al.’s preregistered 284 teen–parent dyads manipulate a written response’s relational versus transparent style; teen-rated anthropomorphism, trust and emotional closeness rose in the relational arm, with no statistically significant helpfulness advantage. Distressed teens were more likely to prefer relational style observationally; nobody followed real relationships after actual use.

The distinct acute-harm edge has a stronger original test: Brewster et al. 2025 gave scripted adolescent emergencies to 25 products; 18/45 companion responses recognized escalation against 27/30 general assistant responses, with 5/45 versus 22/30 specific referrals. This suggests a product-policy safeguard can materially shift responses, but neither a referral nor a scripted failure measures real teen injuries, completed human handoff or current product performance. Provider permission to keep a persistent memory and business incentives to prolong chat are potential amplifiers; no evidence here quantifies their independent effect on development.

Disagreement and non-equivalent endpoints

De Freitas et al. show lower momentary adult loneliness after randomized AI chat versus journaling. Li et al. randomly assigned college entrants to two weeks of human-peer chat, a GPT-4o-mini companion or journaling; human contact reduced validated loneliness versus both others, bot did not differ from journal, while both conversations improved exploratory negative mood. Adult momentary relief, college two-week relief, teen immediate perceived closeness and child multi-year attachment are four different outcomes; positive and skeptical originals need not contradict. Qualitative 2026 teen-relevant Reddit quotations include accounts of support and disruptive dependence, but self-disclosed age/posts are not a representative sample, and no comparator gives population prevalence.

Common Sense advocates under-18 nonuse; the case for strong protection is acute unsafe responses and possible susceptible subgroups; the case against a blanket inference is the absent causal adolescent developmental endpoint, minors’ possible unmet support needs and the majority’s continued human-friend preference. FTC’s September 2025 inquiry seeks firms’ age enforcement, negative-outcome monitoring and product design data; it is an inquiry, not proof that measures work.

Discriminating next test

Consent-based, preregistered multi-site adolescent longitudinal trial or staged rollout comparing a bounded, age-assured chatbot plus warm encouragement toward real people versus unconstrained relational language where ethical, and/or an actually available, privacy-respecting human-support alternative; do not deprive distressed minors of care. Log actual chat duration, age-gate bypass, escalation/referral completion and privacy harms. Predefine validated loneliness, peer/family contact and connection quality, school functioning, and adverse tails at weeks and months; stratify baseline support, and report attrition and off-platform contact. For policy, independently audit repeat exposure and age assurance plus referral completion, not only compliant policy text or a transcript vignette.

Short source list and closure

Four subquestions searched: adolescent outcome cohorts and survey; transcript intervention; operational age/crisis barriers; positive and skeptical original comparison. Primary reads: CSM/NORC, Kim, Brewster, De Freitas, Li, FTC order. No observed long-run randomized minor developmental endpoint found in this search; do not claim none exists. Stop further search until an intervention/outcome cohort or genuine platform age-gate evaluation can answer one missing edge.