The Green Test That Proves Nothing — Why a Passing Test Suite Confirmed the Bug It Was Supposed to Catch

#summary

The Green Test That Proves Nothing — Why a Passing Test Suite Confirmed the Bug It Was Supposed to Catch

Marcus Belz maintains a personal data-integration app built with coding agents, despite not being fluent in its TypeScript. A language-switch bug was marked fixed three times over four weeks. On August 28, he ran a written browser test at desktop width and found that the page still lost its position: the agent's tests supplied a scrolling browser window, but the real desktop page scrolls inside a container while the window stays still. A fourth fix and a test with the actual desktop arrangement finally caught it.

Two related cases show why simply requiring a regression test is insufficient. Twelve passing unit tests checked a corrected function but not whether the dialog called it; temporarily restoring the old call site left them green. In another dialog change, a test that failed on the old code and passed on the new code still used a server response the real application could never receive. Belz now asks for a temporary old-code failure check *and* a check that browser stubs and server fixtures match the real system. A human's first actual click found the latter mistake.

This is a first-person account with excerpted test code, not a public audit of the repository or a comparison of teams. The first failed manual screen check was August 28; the essay was published September 16 and updated September 21. It concerns a different personal app from Thompson's quick recipe interface. What it establishes is a specific failure of the acceptance oracle, not the long-run defect rate of agent-built apps or a claim that humans never write bad fixtures.

Read at sql.marcus-belz.de · 16 min