Belz’s private DI² app: green tests, wrong world, and the manual oracle
Belz’s private DI² app: green tests, wrong world, and the manual oracle
Thing read / dates. Marcus Belz, “The Green Test That Proves Nothing”, published September 16 and updated September 21, 2026; cross-checked against his September 6/18 six-month operator account. DI² is his private Next.js/PostgreSQL ETL-generator project, where he says Claude Code writes TypeScript he cannot read fluently. These are first-person reports with excerpted test code and dated observation, not a public repository audit or agent transcript.
One behavior survived three nominal fixes. A bilingual site's language switch should preserve scroll position. Two late-July fixes and a late-August follow-up all had green tests and were marked done. On August 28, Belz ran a newly written seven-step manual browser case at desktop width and saw the original problem. The agent's test stub offered window.scrollY a nonzero value; on desktop, an inner container actually scrolls and window scroll remains zero. The fourth fix read that container first, then fell back to the window. A replacement test with window at zero and container at 640 failed on the old implementation. A narrow/mobile browser had masked the error.
Two more same-week tests of the test. In a column-validation dialog, twelve unit tests of a corrected pure function passed while the dialog's actual call site still had the old decision: temporarily restoring the old call-site behavior left all twelve green, whereas a new clicked-dialog test turned red. In another dialog change, even a red-on-old / green-on-new counter-check initially succeeded against a fabricated server fixture: its empty affected-column list could not occur in the real request, which always includes the requested column. The first actual click triggered the guard; a corrected fixture and old-guard check then turned red. The first route asks whether the stub matches a real browser; the second whether the tested function is wired in; the third whether the fixture matches the server contract. Belz says five/six later test suites refer back to the call-site case as of September 16/18, but raw private histories are unavailable.
Workflow change and limit. He adopted an explicit temporary regression counter-check (red on the broken code; green after restoring the fix). His reported informal before/after commit count—roughly zero of a hundred recent fixes mentioning such a counter-check before August 28 and nearly half in the next two weeks—is not an independently verified quality improvement. A red-on-old check proves sensitivity to changed code, not fidelity of the fixture to production. Manual UI steps or a real-contract check supply a separate observation. His other September 7 shared-tree account describes five late-August/early-September conflicts, including a 220-process load test that the agent disclosed only afterward. Separate worktrees and a durable bug file helped, but do not make the acceptance oracle trustworthy.
Why here. This is a maintained personal app, not Thompson's five-minute private UI or an audited production team. It puts a price on the promise that a coordinator can summarize agent progress: a green signal can preserve a false environmental assumption for four weeks. A human who cannot read every generated file needs two separable gates: a check that fails on the old behavior and a check that the test-world matches the actual client/server. Compare FastyBird's open hardware ledger and ParallelPilot's status-versus-steering distinction. Missing: full commits/session traces, human-hours, model cost, outside-user outcomes and long-run failure rate.