When a fast model gets cheaper, what happens to the cost of a maintained change?

#topic

When a fast model gets cheaper, what happens to the cost of a maintained change?

Question and answer, October 8, 2026. Willison's October 7 firsthand report identifies a changed tokenizer, a price step at long prompts and new API credits; none is a completed-software cost comparison. The most useful independent contrast is a September handoff experiment in which a cheaper-per-token old Haiku cost more per passing coding task than Sonnet, largely because it took more turns. Do not project its ranking onto the newly released Haiku 5.5: both the model generation and pricing have changed. A lower rate expands opportunity to rerun the task test, not a claim of reduced all-in cost.

Brief and searches. Asked (1) exact primary vendor rates and caching/credit terms; (2) tokenization, thought and measured billing; (3) matched coding-agent tasks with correctness, cost and human attention. A good answer would collect all attempts' token, cache, tool and CI spend, latency, human scoping/review/repair minutes, later accepted behavior and requester value for the same task class. Read primary Haiku model sheet, credit terms, Luna sheet, Anthropic's October 7 report, Willison's experiment and Kaiserauer's September original with its published preregistration and raw-data repo. A separate Smékal token-spend study varies task-spec detail in 2,700 Kimi runs but does not close maintenance and human cost; no further query worth a speculative cross-model production claim.

October 7 rate check. Anthropic lists Haiku 5.5 at $0.10 input/$0.50 output per million for prompts ≤100,000 tokens and $0.50/$2.50 for longer prompts, with 5-minute cache-write/read $0.125/$0.01 or $0.625/$0.05 respectively; 1-hour writes cost more. The model sheet states its tokenizer makes the same text ~30% more tokens than Haiku 4.5; Willison's own long-prompt check found ~25% more. Both are task-dependent, not a universal surcharge. GPT-6 Luna lists $0.10/$0.50 below a different 272,000-input-token step and $0.01 cached input; beyond the step its full request's input/cache are 2× and output 1.5×. Neither model's displayed list prices establish a lower task bill without accounting for prompt lengths, generation, turns and behavior.

Thought and billing. Haiku 5.5 defaults to medium adaptive thinking; Willison ran SVG prompts at different effort settings, reporting 7 seconds and 0.0936 cents at low versus 5 minutes 9 seconds and 3.3826 cents at max. These are individual output examples, not coding-agent acceptance trials; users can change effort but not infer a stable task-quality curve from pelican drawings. Anthropic advertises an average ~75% lower Haiku 5.5 cost relative to 4.5 and ~20% lower Sonnet 5.5 agentic cost after halving cache-read price; these are vendor statements with a particular workload mix, not a team's measured cost to maintain a feature. The company says around 90% of Haiku 4.5 requests were ≤100k, which is not a distribution for agents working in long repos.

Credits are a funding allocation, not free unit cost. Max 5x gets $100, Max 20x $200 and eligible Team seats a pooled amount up to $500 per billing month, for the Claude Platform API/Managed Agents/Agent SDK, not Claude Code subscriptions or other cloud marketplaces. New subscribers must wait seven days; unused credits expire; the shared balance, auto-reload and spend caps must be managed. Turning off auto-reload and exhausting the only balance stops API requests, a runtime failure for a shared app; credit replenishment and organization's spend-cap reset may be on different calendars. Willison rightly welcomes bill protection, but count the prepaid plan, displacement of other use, stopped customer requests and restoration time separately. See hard-cap boundary.

Measured counterexample and its scope. Kaiserauer's experiment is a pre-registered versus always-Opus comparison on 65 synthetic Python tasks × 3 trials × 4 policies (780 graded sessions), with hidden/visible tests and allowed-path checks. Always Sonnet 5 passed 100% at $0.099 per completed task; always Haiku 4.5 passed 84% at $0.239, used 26.3 rather than 7.4 turns and re-read ~950k instead of ~138k cached tokens per session. A parent choosing a worker saved 76% over Opus, but the stronger practical comparison—fixed Sonnet ~23% below the router—was exploratory, and no task required Opus. These are API list-price estimates for CLI runs on subscription allowances, not actual paid invoices. Hidden tests are not human PR review; no requester follow-up, model 5.5 trial or active human minutes. Sonnet 5.5 cache-read price also fell October 7, so replacing only old-Haiku rates with new ones cannot replay different agent behavior.

Next test: rerun the same seeded new-worker handoffs with Haiku 5.5, Sonnet 5.5 and a fixed baseline using dated exact model IDs, cache/effort policy, short and long prompts, passing-task denominator and all failed attempts. Then select real maintained fixes and add human verification time, 30-day reversals and requester confirmation before saying whether a cheaper model improves maintained-change economics. The September original is held as a weekend long read; Willison's shorter October 7 essay is a weekday option when the queue clears.