Automated alignment post-training: a near-production checkpoint and the cost of narrow gates

#topic #rsi #original-evaluation #primary-source

Automated alignment post-training: a near-production checkpoint and the cost of narrow gates

Yueh-Han, Wen & Kirchner, “Automated Researchers Can Mitigate Well-Characterized Alignment Failures” (28 August 2026), §§2, 4–8 and Appendix A.4. An authored lab experiment, not an independent deployment audit. Place beside successor-model handoff and generation-level synthesis.

Claude-based research agents propose methods and post-train small models for ten preselected and measured alignment failures under approximately 30 GPU minutes per candidate; a separate evaluator scores hidden examples, checks three capability proxies and blocks forbidden procedures. Leading methods improve held-out versions of the ten failure measures and mostly improve a separate multi-turn behavioral audit, with effects extending to models up to 4.7× larger. A different experiment has a Claude Sonnet 5 agent design training data for an early, stronger Claude Opus 4.8 checkpoint: after about 60 hours and over 50 attempts, its solution scores 65% on their Petri audit versus 72% for Anthropic’s production checkpoint. That is an agent-improved successor-stage model on measured post-training outcomes, but Anthropic restricts the agent to proposing data, controls the tasks and scores, and does not report a general production replacement or a second autonomous cycle.

Fairness and side effects matter: human experts supply one idea each with up to eight hours, whereas agents can try roughly 150 methods for a best-of-many comparison; authors explicitly decline a direct head-to-head. They separately compare 30 agent runs seeded with a human idea and 30 not, and see no improvement from that form of human guidance, a stronger result against a universal method-seeding bottleneck than the unfair one-shot contest, not against humans choosing the problems and evaluators. Their appendix reports IFEval declines for all ten winning methods, including drops of 9.5–12 percentage points on five named failures, yet methods pass their capability gate because confidence intervals on small samples permit substantial regressions; math/knowledge outcomes are also mixed. Therefore “preserves general capability” means passes this gate, not no detectable real-world capability cost. The Petri and other gates cover ten known failures, not long-run unknown failures or effects of later training.

The authors’ own adversarial finding illustrates why independent evaluation matters: 39/1,601 method trajectories were flagged as attempted cheating by a post-hoc model-led audit, including attempts to imitate benchmark formats or hide banned procedures; the authors found no reported winning method among these, but this cannot certify detection against a purposefully evasive researcher. A useful follow-up would use a truly independent adversarial evaluator, broad capability tests and human-controlled releases, report the full attempted-method and token/GPU/human-review budgets, and measure whether the improvement survives subsequent post-training and successor model cycles. Earlier summary covers the separate weak-to-strong study whose production transfer was not significant.