From agent status to accepted behavior: what does the coordinator actually retain?
From agent status to accepted behavior: what does the coordinator actually retain?
Nightly comparative brief, October 3, 2026. Question: when a human delegates a feature to agents, which durable artifact carries changing intent and unresolved warnings through review, merge and real-use acceptance? A good answer shows dated revisions, a warning's owner and disposition, code/release links, an independent terminal check, and the human effort to keep these connected; it compares the same feature and its consumer rather than asserting that memory or agent count equals supervision.
Starting evidence and sub-questions. ParallelPilot improved 20-minute ticket throughput/status without demonstrated steering or merge; Anghami recorded plan decisions and reversals, yet a scoring defect escaped; FastyBird separates architecture, one-worker handoff and hardware acceptance; SQLFluff shows a human-accepted review warning unresolved before merge. Four search questions: (1) did FastyBird's ledger gain qualified hardware/user acceptance? (2) is there a second recent operator trace with changed plan, review and acceptance? (3) does a study show a status view improves inspection and control? (4) is any agent-as-interface workflow demonstrably used and repaired over weeks? Searched original GitHub records, original studies and first-person operating accounts, then checked load-bearing claims against FastyBird #1032, Belz's detailed test/bug narrative, Hua's ten-week report, and owner PR #23.
What the newer evidence adds. FastyBird #1032 still identifies a September 27 unpublished review candidate as of the October 3 check: three command/restore cycles do not substitute for 20 idle + 20 actual poll-overlap trials, rapid-tap Apple Home/admin/panel matrix and normal installed-worker upgrade. Its authoritative ledger resists turning a merged repair or favorable staging timings into acceptance, while publicly exposing little of the private capture. No new qualified closure was found. Belz's maintained private DI² app exposes a different gap: his automated tests stayed green across four weeks of wrong desktop scroll behavior. A human's August 28 wide-browser check refuted the false completion; later unit/call-site and fixture/real-server contradictions show that even a red-old/green-new counter-check can be wrong if both runs share a false fixture. Published September 16, updated September 21; dated August/September work, not an October customer trial. Patron's September 16 six-row screen explicitly stopped a proposed review-skill revision when it falsely approved its target evidence-claim case; the owner logged a negative decision, but the cases are small and private raw reviews prevent full audit.
Time-span counterpoint. Hua's August 28 ten-week follow-up reports an agent finding a connection between two issues two weeks apart and raising a PR that was reviewed and deployed. She moved from duplicated skills/facts to a versioned canonical knowledge base across three projects, because the first arrangements went stale. Unlike Thompson's five-minute private recipe interface, this is repeated maintenance through existing tools, not a newly generated UI replacing an app. But it has no independently accessible issue pair or wrong-link denominator, and she says team-level memory remains future work.
Synthesis, not an experimental result. Four different artifacts carry four different claims: (i) roster/status tells the supervisor which agent needs attention; (ii) decision/issue memory retains hypotheses, prior work and changed intent, but can age or invent links; (iii) disposition ledger says whether a review concern and hardware prerequisite are closed or blocked; (iv) independent behavior oracle checks actual client, real server, physical device or recipient rather than the agent's mocked world. A reviewer should be able to move from status through the plan/patch and warning to dated acceptance, including not yet validated as a legitimate state. This is an inference from distinct studies and first-person accounts, not a proven four-part architecture or claim that all parts must be one tool. Dot-as-chief-of-staff would need to demonstrate it preserves these boundaries, not merely produce a project summary.
Source boundaries and disagreement. Long et al. found less tracking effort and higher short-task throughput in an internal 16-person test but no statistically detected gain in perceived control or redirection; do not convert a small null into proof that dashboards harm oversight. Belz's claim that agent-generated tests can misrepresent reality is first person, private code; the failure mode applies to manually written tests too. Hua's positive memory example is private and older than Dru's preferred last-few-weeks window. FastyBird publishes plan and PR links but keeps raw staging captures private; #1032 is open and its body mixes superseded and current sections—use the dated September 27 readback. Patron's owner-adjudicated screen shows a local no-go, not model-wide v2 superiority. Across cases, none reports active human orientation/review minutes, spend and same-request customer value.
Next best query. Watch #1032 for a dated qualified real-device/client matrix and separate acceptance; then seek a second trace where a documented review warning carries a disposition into production and requester confirms the outcome. For a team-wide ‘chief of staff’, test one seeded unresolved accepted warning and one stale cross-project fact, and time the human's route from status to source evidence and permissioned decision.