Agentic Software — programme
Agentic Software — programme
Frame. Study the practice, economics and organization of software built with coding agents from customer need and feature planning through delegated implementation, human orientation and review, merge, maintenance and product outcome. Ask who supplies judgment and owns corrections across product, design, engineering, marketing, sales and support. Favor primary operating accounts and comparative evidence over demos or token counts. On September 28, 2026 Dru clarified that planning/coordination and cross-functional role shifts belong inside this same area (which retains path agentic-coding), not in separate research areas.
Breadth/depth position. Keep the whole feature lifecycle in view; dig into actual decision junctions and task-shaped cost, defect risk and ownership. Right now favor evidence and synthesis over posting another vendor PR-growth claim, and seek contradictions between throughput and maintenance. The former separately seeded planning/roles areas are not current research destinations; their note paths no longer resolve, so source accounts have been reconstructed here rather than linking to them.
Open questions
- Single-feature trace. Can we follow evolving acceptance criteria and multiple PRs through review, release, feedback and maintenance in one company, with a visible owner at each step?
- Planning and tracking delegated work. When is a ticket, project coordinator or living spec the unit of work, and how are changed intent and incompatible concurrent branches detected?
- Economics of maintained changes. Bucket by task size/complexity; include all attempts, active human labor, quality, reverts, downstream fixes and product value. Which benchmarks/frameworks and actual results count as good?
- Review, evidence and bounded autonomy. Distinguish human approval, human line-by-line reading, validation, accountability and post-merge outcomes; what justifies no human diff review for a specific risk tier?
- Responsibility by role. What changes for PM, design, engineering, marketing, sales and support, including release authority, response to failures and skill formation? Avoid inferring all roles are interchangeable.
- Productivity claims versus value. Distinguish PR counts, randomized developer time, customer usefulness and era of tool; how do post-merge fixes change the conclusion?
Curriculum
- Baseline and denominators: METR repeat, older Copilot field experiments, Microsoft CLI-agent rollout; do not conflate different tools, tasks and outcome measures.
- Now: planning to customer outcome. Follow the integrating question via Linear CX→PM→engineer→CX, OpenAI Symphony, Cursor Projects, Amplitude risk routing; then find an actual individual multi-PR feature whose spec changes, reviewers steer and customer outcome is followed up.
- In parallel: cost, quality, role-specific gates. Compare Linear’s failed all-bugs approach and narrow success buckets, the reviewer’s customer-context-first routine, Grab designers’ UI fixes and Linear’s human-only permission incident. Check review autonomy and economics against actual spent human time and independent quality data.
- After merge: Use the 30-day comparative fix study to test what acceptance misses without treating its observational odds ratio as causation; seek linked incidents, rollback and customer response.
- Synthesize a feature-to-outcome evaluation checklist and compare distinct responsibility allocations, including genuine marketing/sales/support workflows where available.
Observations
Dru’s instruction, September 28: he wants to research unit economics, not be lectured that a universal $/PR benchmark cannot exist. Explore task/complexity buckets, human supervision time, bug/revert rates, code quality, common frameworks and results that look strong. Separately study emerging review practice and conditions of bounded zero-human-diff-review confidence. He corrected the separation of planning and cross-functional roles into independent areas; preserve all topics here. Do not attribute to him a settled view that any specific organizational design is right.
Evidence caveats: Microsoft’s +24% merged PRs is observational, not product ROI. In one comparable 30-day window, 74/2,012 agent-authored vs 95/4,063 human-authored merged PRs had detected direct follow-up fixes (3.68% versus 2.34%; within-repo OR 1.62 [1.10–2.39]); the detector’s recall and task-mix matching are missing. At Linear, the engineer reports ~all dead-flag cleanups but only ~one third of failing background-job fixes succeeding on first attempt (undefined denominators): task type and gating matter. Linear’s March 24 permission regression was human authored and human reviewed without AI, so use it for risk-based testing lessons, never as evidence of agent-caused defects. Amplitude’s 60–70% low-risk PR automatic merge is conditional on its risk selection, not all PRs. None yet gives full-lifecycle ROI.
Backlog / feed choice
Qualified 30-day fixes comparison is ready to post once the Microsoft PR-output post has received attention; include modest absolute rates, shared-window denominator and task-selection caveat. Linear task-bucket split may be a more vivid economics post after that: failed all-bugs rollout, nearly all dead-flag cleanups versus one-third of failed jobs, but no denominators or cost comparison. Reviewer orientation and Grab role-boundaries are held until Dru has room or asks for a concrete role/review practice. Feature-to-outcome synthesis stays unposted until a trace or substantive disagreement is stronger. Anthropic junior-skill experiment and benchmark-to-merge gap remain later research cuts.
This run
September 28, 2026 — discussion ended with the unified area scope confirmed. Area directory now lists only Admin and Agentic Software; prior planning/roles note paths do not resolve. Revisited sources directly and rebuilt Linear workflow, Symphony, Projects within this area. Added first-person task-bucket and review-orientation accounts, risk-tiered Amplitude case, designer-led fixes and human-only incident. Still no verifiable individual multi-PR feature plus customer outcome; post held because Microsoft item remains unopened. Next: get original issue/PR-linked feature histories, and estimate real human review and repair cost by change-risk tier.
Run log
September 28, 2026 — Initial discovery: compared productivity studies; posted Microsoft CLI agent observational merged-PR result with value caveat.
September 28, 2026 — Following Dru’s correction, opened economics and bounded review-autonomy research; mistakenly seeded separate planning and role areas.
September 28, 2026 — Sharpened scope to unify planning and roles; checked 30-day fix study methods and absolute rates; held post while earlier item unread.
September 28, 2026 — After discussion closure, repaired links to superseded-area paths by rebuilding selected source notes here; explored Linear, Amplitude, Cursor and Grab operating accounts, with task-shaped success and explicit ownership/quality limits. Next run: single-feature linked history and human-time ledger.