Voice phishing survey: cheap scalable callers, intentions not transfers
Voice phishing survey: cheap scalable callers, intentions not transfers
Fred Heiding et al., “Evaluating AI Models’ Capability to Automate Voice Phishing Attacks,” July 2026 preprint. The file carries a journal volume dated 2027; as of 5 October 2026 describe it as a preprint, not a completed future-year publication. Fraud chain | Email trial.
YouGov recruited an opt-in US adult panel, weighted/matched 4,100 internet users to demographic distributions; each saw one assigned recording or transcript from five scam scenarios, with AI voice/script and human voice/script comparisons. A supposedly familiar sister’s cloned voice was not the listener’s actual relative. Staff edited the demonstration recordings to remove model refusals, AI-identifying beeps and disclaimers: a best-case clean excerpt, not a successful uninterrupted attacker call. The reported 16.5% pooled willingness includes yes or unsure answers, not payment or credential submission; up to 36.1% in the relative-distress condition likewise cannot stand for funds stolen.
Some synthesized voices rated comparably to human voices, but quality gaps across models did not significantly change reported compliance in the authors’ model comparison. Their profitability table assumes a rate at which persuasion becomes payment, a mean receipt per victim and US human labor cost. It yields ~$1–$3 per hour for three AI voice systems, negative for others and a US-wage human, not observed crime revenue; criminal labor markets need not pay US wages. The source does show a plausible marginal labor-cost reduction given suitable voice API and delivery channel, while provider safeguards, phone delivery, target authentication, financial transfer and downstream freezes remain untested. Needed next: an ethical live-stakes transfer proxy, balanced human operators at actual wages, unedited safeguards and tracing from attempted calls to bank-approved and unrecovered transfers.