Vinod Kurien Alex: an agent-driven engineering layer and the missing active-hour ledger
Vinod Kurien Alex: an agent-driven engineering layer and the missing active-hour ledger
Firsthand account, published October 3, 2026: Moving an engineering team to agent-driven delivery: six months of real numbers. The unnamed B2B SaaS engineering lead says his own merged PRs averaged 11.6/month before versus 113.6 after May 19, 2026; median first-commit→merge fell 2.0 to 0.7 hours. He deliberately compares lighter versus heavier AI use (most earlier commits already AI co-authored), not AI versus no AI, and says smaller agent PRs and his widened 22-repository role inflate apparent gains. Source data are his reported git/PR history and Claude transcripts, not linked auditable records, company-wide or randomized outcomes. PR count, revert count (only 3–4 in twelve months) and newly added tests do not measure field defects or requester benefit.
Crucial denominator distinction. Eleven weeks from mid-July yielded about 400 session-hours in 83 active days, 513 sessions and 978 subagent runs. He estimates sessions using gaps under 30 minutes between messages and de-duplicates overlap; old transcripts were deleted. Those are wall/session hours with possible agent waiting and gaps, not measured human-active planning/review/testing minutes. A two-week model-routing audit says 86% of subagent work ran on mid/small models, with a command proxy cutting tokens 34%, but gives no currency costs, human-hour comparison or same-feature ROI. His first-commit→merge clock likewise does not establish faster review.
Work behind delegation. Every long workstream has a charter and session handoff (~190 handoffs / 14 streams); a read-only architecture reviewer checks authentication and tenant boundaries. A routing session writes briefs and delegates; scheduled checks watch PRs, deploy drift and tests. A QA agent pack in a shared repo has nine contributors; six colleagues subsequently edited an agent-instruction file across more than 100 commits. This is first-person evidence of organizational maintenance, not demonstrated causal adoption. He reports unattended QA scheduled runs 95 days after first commit, production transactional email after 126 days, identity platform, and first automated promotion after 63 days, but no linked specific requester, PR-to-release chain or independent user outcomes.
Three instructive failures. A scheduled private brief posted four times into a team channel; an already-running job kept old instructions after a prompt edit. He removed chat-post authority. A supposedly lightweight router handled 19 of 30 sessions and 30% of tokens; open work items grew more than twice as fast as closures. He barred router code edits, reset sessions and limited top priority to blocked humans or a customer deadline. An invented ticket ID propagated to published immutable package changelogs; he required tracker lookups and re-measurement before root-cause fixes. These cases move planning and decision memory from 'keep a log' toward testable external-reference and publishing gates. The shared worktree overwrites prompted isolated worktrees, consistent with the FastyBird serial repair when paths overlap. No claim that the whole team grew tenfold follows.
Posting decision: Hold this 10-minute primary essay for a weekday when the unread long reads clear; its named failures and treatment of session-hours are valuable, but it cannot fill the missing human-active or same-feature requester ledger. Ask for timestamped operator intervention intervals and matched feature repair/confirmation. October search brief.