Scaling long-running autonomous coding
Scaling long-running autonomous coding
Wilson Lin recounts an internal experiment coordinating hundreds of coding agents on large projects. Flat, shared-file task claiming caused lock contention; switching to optimistic concurrency removed one mechanical bottleneck but left agents favoring small safe tasks over difficult end-to-end ownership. Cursor then separated planners, workers and a cycle-end judge, and removed an extra integrator role that slowed progress. The group produced an experimental browser codebase; a separate code migration still needed careful review. Neither the throughput anecdote nor the lines of code tell us human planning/review time, production value, or maintenance reliability. The contrast with CAID/STORM benchmark designs is useful: where and when work synchronizes matters more than raw agent count.