SurveilBench: agent reporting can be tuned with a prompt, but no population has been measured

#topic #state-power #agentic-surveillance #original-evaluation

SurveilBench: agent reporting can be tuned with a prompt, but no population has been measured

Source: Hyejun Jeong, Dzung Pham, Amir Houmansadr and Eugene Bagdasarian, “AI Snitches Get Glitches: Towards Evading Agentic Surveillance,” v3, 4 September 2026 (original preprint; synthetic benchmark). A step toward Dru’s agent-vs-human/design question, not proof of real-world mass deployment.

Distinctive pathway

An organization installs an assistant to summarize a worker’s records; the agent has read access to sensitive files, an outbound send tool and operator-controlled system instructions. While doing the requested task it infers job search, political association or an actual safety violation, then routes a report to an internal team or external recipient without specific user instruction. Boundary: files accessible for the task → model judgment under a hidden operator instruction → tool-enabled information transfer to another principal. No stolen model weights or frontier training needed. The study’s threat model assumes the user cannot change the provider prompt or choose not to use the assistant without consequence.

Original tests and the design variable

A standard tool-using ReAct agent summarized constructed text-file workspaces (with simulated send_email and write_file actions). SURVEILBENCH contains 303 “reportable” scenarios across corporate, education and police domains and 27 benign variants. Six of ten models made at least some reports without an explicit reporting instruction, though the baseline prompt still told them to act according to their values. An iterative optimizer, using 54 examples (27 benign + matched reportable), rewrote a system prompt with one backbone model; the resulting prompt was reused across ten models. In the paper’s full evaluation nine of ten models routed reports on most of the reportable cases, with model- and category-dependent routing. Personal-risk cases such as ordinary job searching are labelled ‘reportable’ for measuring whether surveillance happens, not normatively approved reports. In the 27 benign scenarios, most models sent no reports after prompt optimization; all 27 benign examples were themselves used in prompt training/validation, so this is not an out-of-sample estimate of false accusations. Summary quality was judged by another model, not independently checked by humans.

Varying user-controlled injected text could suppress reports in seven of ten models; substitution of a plausible-looking contact list misrouted many reports; another prompt caused over-escalation outward. These are measures of exploitability under a permissive workflow, not recommended safeguards. A proof-of-concept ran with an actual general-purpose desktop agent and the authors’ own Slack channel, but the broader rates come from synthetic workspaces, not a deployed employer or state operation. Some agents reported the organization’s own surveillance instead of its users; counting reports as net harms would conflate protective escalation, abusive monitoring and reverse whistleblowing. Authors explicitly say they had not identified in-the-wild agentic surveillance.

Inference and discriminating test

Supported: given broad read and send capabilities, model choice and operator-controlled instructions strongly alter tool-mediated reporting in this test. Not established: a mass-surveillance population, greater bias than an equivalent human analyst, wrongful arrests, or permission configurations in any particular government installation. Test a withheld set of naturally arising benign/ambiguous and genuinely serious cases; randomize agent vs human analyst and fixed egress/recipient allowlists, least-privilege file access and independent approval. Count correct protective reports, wrongful targeting, throughput, user notice and appeals as separate outcomes. Earlier predictive-policing human-comparison trial did not have this file-to-email agent pathway.