When an agent builds an app, who owns its runaway bill?
When an agent builds an app, who owns its runaway bill?
Question and research window, October 4, 2026. Does a provider's hard cap turn a disposable agent-built interface into a bounded ongoing liability, and what happens to users when it trips? Good evidence names the billable service and authority, cap's rollout and exclusions, the stopping event, recovery owner and user outcome. This is the deployed application's runtime expense, not merely the coding agent's inference quota. Started from Thompson's five-minute personal UI, maintained-change economics and independent acceptance. Subquestions searched October 4: AWS rollout and scope; Google Cloud eligible services and stop/restart behavior; a firsthand agent-app billing failure; and the failure state of capped services. Primary sources below supply the answer; no repeat queries needed for now.
A new argument, not an incident report. In his October 3 essay, Simon Willison argues usage-based providers should default to a cutoff returning errors at a monthly limit, with an explicit opt-out for those who prioritize continuity over a surprise bill. His examples concern agents lowering the friction of creating apps that call paid services; he does not report testing either new provider cap or tracking cost for one generated application. The operator should choose the tradeoff before the code runs, and the agent could warn against deploying on uncapped services. Provider docs qualify his shorthand.
AWS as of October 4. AWS Account Management documentation says the new experience is rolling out to a limited number of customers; a paid-plan project owner can set a project-level, pre-tax monthly limit, minimum the greater of $20 or a usage-based estimate. At the ceiling AWS pauses the whole project and stops resources, preserving data; the owner must increase the limit to reactivate, and some resources need manual restart. Without action, paused project data is permanently deleted after 90 days. The docs expressly frame it for experiments and say production use fits only when a brief pause is acceptable. Optional earlier controls can block new launches or pause idle/top-cost drivers before the cap. The September 16 announcement promises this experience; neither launch nor Willison's essay establishes availability for ordinary existing accounts or that it is on by default.
Google Cloud as of its September 30 documentation update. The spend-cap guide labels the capability Preview and limits each monthly cap to one eligible service in one project: Gemini API, Agent Platform, Cloud Run or Cloud Run functions. A triggered cap blocks new usage for that service, not other services or ongoing fixed compute/storage charges; in-flight work completes, and reporting/enforcement latency can leave billable overages. A billing admin or authorized owner must manually lift the cap; affected services can take up to an hour to recover. If lifted in the same month, the cap will not retrigger unless the amount is increased. July 28 launch post confirms public preview and single-project/service scope. Neither vendor's new controls imply a universal default, third-party API protection or uninterrupted service.
Firsthand incident, but different causal pathway. In an August 31 METR postmortem, a researcher deployed an agent-orchestration dashboard on a personal EC2 instance behind intended Google authentication. A fail-open flaw exposed it for days in March; attackers prompted an agent to disclose a public-model API key and used it for three weeks. METR estimates roughly $600,000 in free granted credits consumed, not a $600,000 cash bill. It says no spend ceiling was available for that key at the time and that noisy rate-limit errors plus an incomplete usage dashboard hid the abuse. METR added public-app security review, narrower controls and spend alerts where possible. This is an unauthorized-attacker/key-exposure case, not a coding agent innocently provisioning too many resources or a test of the later AWS and Google caps. It shows how generated app security, access scope and usage monitoring interact with the cost boundary.
Synthesis and disagreement. Willison prefers default financial fail-closed over a surprise bill; both provider docs warn that enforcement can disable real service, and Google Cloud explicitly reserves latency/overage and fixed charges. Neither alternative is a free default for a critical shared service. Proposed release/operations record: owner-approved cap and scoped identities; expected cost and alert recipients; for each triggered cap, affected customer requests and failed actions, data-retention clock, authorization to lift, restart check and requester confirmation. This is an inference from the provider rules and METR incident, not a measured reliability improvement. For private experiments a service pause may be tolerable; for a customer-facing workflow measure user loss from pause alongside surprise-spend exposure and human recovery time. Still missing: a same-app comparison over a month of generated app build cost, actual metered runtime charges, cap activations, customer impact and human recovery. The METR episode concerns March work published August 31, not a last-few-weeks deployment.