Rover’s Agentic Afternoons: an afternoon of practice before a merged PR
Rover’s Agentic Afternoons: an afternoon of practice before a merged PR
Original account: Nadia Wadzinski (senior staff engineer) and Riccardo Canalicchio (engineering manager), Rover engineering, November 1, 2025. An operational account, not an experiment with a control arm despite its authors calling the workshop an experiment.
What actually happened. Developer Experience first tested agent tools and prepared a centrally configured Codespaces installation, a shared CLAUDE.md for app layout and build commands, adoption/cost hooks, and an educational talk about how agents work. They then reserved a company-wide afternoon for hands-on Claude Code on participants’ own small backlog tickets, lunch and instructors in Seattle and Barcelona. Engineers were asked to generate a PR without hand-writing code as a learning constraint, while planning, steering and reviewing every generated line; authors explicitly say it was not normal-workflow policy. An initial agent plan reviewed and changed before execution and careful review formed the pause points. A shared task log tracked which work succeeded and whether manual intervention after the workshop was required to produce a mergeable PR, but its rows and human minutes are not published.
Denominators and aftermath. Three-quarters of engineers attended and tried the tool; about half of participants reached a merged PR. The authors say weekly active coding-agent use rose from under 20% before to over 50% the week after; at writing, 77% of engineers were active users. No number of engineers or merged PRs, quality, downstream value, total teaching/setup time, or comparison group are reported. Small achievable tasks were intentionally selected, so half merging is not a completion rate for arbitrary backlog. They report lunch spend slightly exceeded LLM token spend at the event but omit both amounts and cannot infer that tokens are cheap relative to the engineering hours of a company-wide afternoon. Later continuing use is suggestive uptake, not proof of productivity or net ROI.
Where it moves the question. This is the concrete missing bucket in a naive cost/PR calculation: instructor preparation, shared environment/documentation, a half-day participant opportunity cost and post-session manual merge work, offset by any resulting skill/reusable setup. Tests for team artifacts and human learning should follow the shared task log into later independent task completion and review quality, not substitute a weekly-active chart. Compare with the controlled unfamiliar-library skill test, which measures understanding but not a full coding-agent organization rollout.