Anghami Mikrowave: versioned decisions across agents, and bugs after review
Anghami Mikrowave: versioned decisions across agents, and bugs after review
Read: Elias El Khoury, “A Year of Building a Media Pipeline with Coding Agents”, September 29, 2026, first-person Anghami engineering account. The practice spans September 2025–September 2026; the most relevant parallel-work records are July–September 2026, including three stacked PRs during the week of September 7. Mikrowave product page describes internal production and September 2026 sample title comparisons, but both pages are from the operator/vendor. The code repository, PRs and session log are not linked for external audit.
Operating loop. The engineer wrote a 333-line specification in September 2025 to replace disposable bash experiments with a Python media service; it carried the failure learned in the prototype into acceptance tests rather than preserving the bash implementation. By March–April 2026 Claude wrote code in a sandbox while Codex reviewed plans. One April plan about scoring encodes underwent five human-adjudicated cross-model review rounds in two days; three successive alignment errors were caught before shipping. By August–September Claude often implemented in parallel worktrees; Codex review was mostly invoked as a tool by Claude: all 33 Codex sessions from August 1 to September 10 in his retained records had no direct user messages. The human still decided which review findings to accept, ran test gates, approved branches and deployed from outside the sandbox; the agents had no credentials to push or deploy. A September workload had 28 workflow calls in one week, 17 on one day, alongside three stacked PRs going through plan, implementation, review and merge. These are activity counts, not time saved.
How intent survives. Plans graduate to repository design proposals, implementation plans, measured policy numbers and resolved-decision documents, with decisions and reversals written for a new session. The account reports 13 such docs (~13,000 lines) and 44 compactions across 42 retained Claude sessions in July–September. CLAUDE.md/AGENTS.md encode past failures, including a heartbeat block and wrong HTTP 409 behavior. Golden tests pin the actual encoder and packager command lines and change only by deliberate decision. This is a counterexample to treating the PR or chat summary as the whole plan: the enduring artifact is a revisable, versioned decision history and its test oracle. It does not establish how a team of humans would share ownership.
Failure and repair. In September, canary titles missed their quality floors; an isolated reproducer showed 7.2% of frames paired with the wrong reference due to timing-precision differences. The engineer added an end-to-end regression test and made a one-line clock-consistency change. This is the same broad scoring feature family as the April review, not proof that the particular reviewed plan or three stacked September PRs caused it. In August a concurrency change triggered intermittent libvmaf crashes; the first agent spun on hypotheses, a freshly briefed second agent traced an upstream library threading defect, and the operator pinned a fixed version. One crash in 46 control runs versus zero in 150 patched runs was too sparse to establish an effect alone; upstream root cause and version floor were the stronger checks. The author explicitly declines to attribute improved code quality to later models: task mix, maturity and his skill changed together.
Integration and cost. A workflow exhausted the operator’s usage limit in minutes assigning expensive models to minor work; a rule now assigns cheap models to mechanical steps and keeps a high-end model at the top. A 13-minute agent item included three minutes of an unnecessary full test suite; the operator took ownership of parallel test gates and asked agents to run only named tests for a small bug fix. The account gives wall/model examples and one person’s decisions, not active human specification, review, merge, debugging or training minutes. Two internal production clients and 63 release tags are reported, but the promised separate product/results account is not yet present; there is no requester confirmation, field error-rate comparison, or counterfactual. This is a stronger recent operating record than a feature demo, but not the full maintained-feature cost/outcome ledger.
Next questions. Can a public artifact or follow-on product account connect one written plan, three stacked PRs, acceptance change, release, later defect and internal user result? What is the human attention budget by stage? Would the golden tests have prevented the September scoring mismatch if the reference clocks had been modeled? Connect to planning and tracking, recent-feature brief and discarding the prototype.