Kaiserauer: routing agent handoffs is not automatically cheaper than one good default
Kaiserauer: routing agent handoffs is not automatically cheaper than one good default
Original, September 2026; read October 8. Andrew Kaiserauer's 16-page white paper asks whether a parent agent that chooses a worker model at handoff reduces dollars per passing coding task. A public harness, tasks and result files and analysis plan make the denominator more inspectable than a list-price projection. Its confirmatory hypothesis is parent-picked worker versus parent-model (Opus) worker, not router versus best fixed worker; author dates the frozen configuration to September 24. The analysis-plan document itself still says its status is draft until a preregistration table is written into configuration, so the report's confirmation depends on that configuration/hashes, not the document's header alone.
Methods and result. Sixty-five purpose-built, clearly specified Python tasks in three small fixture repos (29 fixes, 13 features and smaller categories), each repeated three times in four routing policies, yielded 780 graded sessions. Forty turns maximum, hidden and visible tests and out-of-scope-file checks; costs estimated at Claude API list rates from Claude Code CLI sessions under subscription login. Parent-picked model passed 100% and averaged $0.129 per passing task, versus 99%/$0.528 for always Opus 5; 76% less on this registered contrast, which is a weak and expensive default. The exploratory fixed Sonnet 5 arm passed 100% at $0.099, ~23% cheaper than parent choice; no task required Opus. Fixed Haiku 4.5 passed 84% at $0.239 per pass though its token rates were half Sonnet's: 26.3 turns and about 950,263 cached tokens re-read per worker session versus 7.4 turns and 137,669 for Sonnet. The matched figure counts unsuccessful attempts; it is not a cost/PR figure.
Why it matters. Price per token, model-induced extra turns, and cache re-reading combine; router needs to beat the best fixed baseline rather than the priciest parent. The author's attempted cascade accepted partly broken work because visible tests did not catch missing hidden behavior; forcing a hidden-test oracle for escalation could leave Sonnet anchored on the weaker worker's partial edits. This echoes Belz's fixture problem: a reliable stop/hand-off oracle is part of economics. The same test suites here are synthetic, not observed maintainer approval. September 24–25 model versions were Sonnet 5, Opus 5 and older Haiku 4.5. The October 7 Haiku 5.5 launch and Sonnet 5.5 cache price require an actual rerun; old ordering cannot be transferred. Briefs were shorter and more uniformly scoped than the author's own real handoffs, and no active human hours, customer outcome or post-release fixes were measured.
Held: white paper is a weekend long read for October 10–11 if the queue clears, preferable to a generic model-price announcement. A future weekday can instead use the author's shorter first-person write-up if latest-model replication or a genuine team case arrives.