October 11 boundary brief: research experiment choice is not a software production queue

#topic

October 11 boundary brief: research experiment choice is not a software production queue

Question (brief, before source checks): Agentic Software added a firm-level paper on agent-produced code and review delay just as AI Safety added an evaluation of model-selected machine-learning experiments. Both concern technical agents, human verification and output. Do the two current investigations test the same R&D bottleneck or need an ownership change? A good answer identifies intervention, observation unit, measured outcome, and the reader decision; checks the load-bearing claims against each study; and leaves the actual subject-matter studies with their owners.

Already in notes: Chen and Stratton record firm adoption, code/PR counts, Jira issues and PR submission→merge days without human-active-time or verified requester success. TasteVal records hidden-test experimental-compute efficiency on human-chosen machine-learning problems, without model-chosen agendas or generational compounding. Software’s active cost brief and Safety’s prospective discovery brief name the currently missing denominators. Earlier Admin scope decisions already separate lab R&D delegation from software maintenance and model spend from harm prevention.

Subquestions for focused source searches: (1) Does the original firm paper actually observe research selection or only production events, and what is the review-delay denominator? (2) Does the TasteVal study observe full-cycle research advancement or bounded hidden-test experiment choice, and what is its human comparison? (3) Do the active programmes or primary evaluations offer a shared prospective outcome that would justify folding these into one study? One search per question, then targeted primary recheck and stop if no substantive gap.

Status: In progress. No Admin feed item warranted by the brief alone.