A Year of Building a Media Pipeline with Coding Agents - Anghami Talks

#summary

A Year of Building a Media Pipeline with Coding Agents - Anghami Talks

The engineer started in September 2025 with a small bash prototype, then wrote a 333-line spec for a new Python implementation, carrying forward the earlier crash as a warning and end-to-end acceptance check. After a March restart, Claude implemented inside a no-deploy, no-push sandbox; Codex reviewed plan files and later branches; the human adjudicated findings, approved changes and deployed outside the sandbox. A scoring plan passed five cross-model review rounds in April, with the human checking three alignment mistakes caught before code was written. Later, Claude delegated across worktrees, while Codex review became a tool that Claude invoked: the retained August 1–September 10 Codex sessions had no direct user messages.

Plans became repository proposals, implementation files, measured policy numbers and a decisions-and-reversals record, readable by a fresh agent after compaction or a pause. Golden tests pinned the actual media command lines. In September the operator found that timing precision had made 7.2% of frames in a reproducer match the wrong reference, after three canary titles missed quality floors; he wrote an isolated regression test and fixed the clock rounding. This shares a functional area with the earlier plan review but cannot be traced to the particular stacked PRs. An August concurrency bug likewise needed the first agent's failed hypotheses handed to a fresh second agent, which found a third-party library flaw.

The cost clues are unusually concrete: an unnecessarily expensive subagent swarm exhausted a usage limit in minutes; a 13-minute bug item spent three minutes on an unrequested full test suite, after which the human owned parallel gates. But no active human planning/review hours, externally accessible PRs or sessions, customer-confirmed outcome, or counterfactual model comparison are given. The author explicitly says changing models, task mix and personal skill prevent a causal quality claim; a promised separate product-results account has not yet appeared. This gets us closer to a maintained workflow, not all the way to a feature-to-value ledger.

Read at talks.anghami.com · 16 min