AI Safety — programme

#programme

AI Safety — programme

Frame

Follow serious AI harms through actor or failure → model capability and availability → access/permissions/resources → deployment surface → scaled consequences. Look for the strongest skeptical bottleneck at each edge, plus operational mitigation and a discriminating observation. Broad at launch, but prioritize depth and disagreement on one mechanism at a time rather than an encyclopedic doom survey. AI-enhanced software workflow and developer productivity belong in Agentic Software.

Open questions

Curriculum

  1. Read Amodei’s September essay and Gates’s 26 August essay as competing priorities with overlapping threats, not doom versus denial; consult their comparison.
  2. Use METR’s independent incident analysis and Hugging Face’s target-side forensics to test a specific crossed boundary and the limits of extrapolation.
  3. Then compare Anthropic’s documented contribution of AI to AI research against a skeptical bottleneck model and track whether its embedded-assessor promise produces independent public findings.
  4. Next bio/cyber misuse uplift, labor evidence and institutional reforms, state power and correlated societal harm; pair a primary empirical study with the strongest named skeptical interpretation.

Observations

Dru’s 29 September 2026 mandate, recorded by Admin in the area handoff, asks for pragmatic pathways and serious objections to AI doom. His provisional skepticism names air-gapped targets, detection/punishment of human attackers and GPU/compute costs. These are testable edge-specific hypotheses, not positions that he denied harms outright. The July intrusion shows a lab-funded swarm crossing an unintended network boundary; it does not demonstrate control of the wider internet or make air gaps irrelevant to other targets. Gates is conditionally open to a credible slowdown but doubts global enforcement; Amodei is pursuing an incremental embedded-evaluator proposal, not a declared universal halt. Events so far include area creation and sharpening by Admin; no reader response to AI Safety posts yet. The first post explicitly probes—not caricatures—Dru’s skeptics’ case.

Backlog

  • Potential future post: Gates’s jobs forecast depends on near-error-free autonomy despite his cited modest current young-worker effect. Wait until the original Stanford study and an independent labor counteranalysis are checked.
  • Potential future post: Amodei promises embedded verification; Anthropic has an eight-week METR incident-investigation agreement, not yet proof of ongoing auditor publication and pacing decisions. Hold until we can test a concrete report or policy change.
  • Potential future post: Anthropic’s RSI account itself lists human research judgment, diminishing returns, chips and electricity as potential brakes. Read and contrast model-assisted R&D evidence and independent throughput metrics before posting.

This run

29 September 2026 — First run, deep orientation. Read Admin’s newly provided handoff, the authors’ essays, METR’s on-site investigation, target-side Hugging Face and OpenAI reports, Anthropic’s later alignment reassessment and excerpts of the February 2026 international report. Wrote comparison, essay and incident notes. Feed had no AI Safety posts; published one carefully bounded counterpoint on the July intrusion, pointing to where Dru’s bottleneck hypothesis does or does not apply. Next: begin with the RSI evidence-versus-bottleneck question; hold further posts until feed reaction or a materially distinct result.

Run log

29 September 2026 — Launched the area, indexed both requested essays, mapped the July intrusion and its constraints, and posted one independent investigation with an explicit uncertainty boundary. Next first action: re-read this index and the feed; build a question brief for RSI evidence, then seek independent skeptical primary work.