Hua’s ten-week agent interface: canonical memory, repeated maintenance and still-human review

#topic

Hua’s ten-week agent interface: canonical memory, repeated maintenance and still-human review

Thing read. Data scientist xiaoye_hua's August 28, 2026 firsthand ten-week follow-up to a June personal git/agent-memory experiment. She manages three ML projects; reports both maintenance (alert → diagnose → patch → deploy → verify) and development (assume → code → evaluate → decide). This is a narrated personal setup across work projects, not public code, a controlled comparison or a measured staff-productivity gain.

What changed over ten weeks. Generic CLAUDE.md instructions missed project-specific metrics and staging deployment; she moved procedures into skills. That solved a workflow problem but produced stale/duplicated facts and separate skills she could no longer inventory. She gave one agent a knowledge base as canonical source for skills, issue records and session notes, with saving/loading built into its standing instructions. In one maintenance example, an alert about a seemingly unrelated issue arrived two weeks after a first issue; she says the agent retrieved the first issue record, diagnosed the connection, proposed a fix, raised a PR and, after review, the team deployed it to production. There is no linked PR or repeat-incident rate and no proof that the agent rather than the human correctly established causality.

Stewardship remains. Three copied project-specific agent definitions drifted. She split shared agent definition, cross-project memory and each project's code into separate version-controlled repositories; generic procedure changes go to the first, project facts to memory. She explicitly leaves a team-wide knowledge base and an automatic ML iterate/evaluate workflow for the future, with human validation still intended. ‘Agent as interface’ here is a repeated-use interface to tools and history, not a disposable generated app; versioning, fact authority, review and deployment have not disappeared. Compare Thompson's single private UI iteration, Dot coordination inference, and Anghami's decision log.

Boundary / next test. Her reported issue linkage is one positive anecdote with non-public records; a project-memory steward should also count false links, forgotten changes, stale fact uses, cross-project/context leakage, review minutes, later production recurrence and whether the requester considered the second issue closed. This is a stronger time-span counterpoint to a five-minute UI, but the account predates Dru's October 2 preference for work from the last few weeks.