Do AI-R&D benchmarks measure research judgment or merely scoreable engineering?
Do AI-R&D benchmarks measure research judgment or merely scoreable engineering?
Opened at Dru’s request in discussion of AI-assisted R&D versus self-sustaining RSI, 30 September 2026, 07:20 UTC. His working reading: the September METR result is mild evidence against present-day RSI takeoff; benchmarks tend to saturate and simplify reality, and this model’s incremental gains have not shown a qualitative gain in taste, foresight or independent judgment. This is an objection to extrapolation, not a claim that AI cannot help research.
What the leading families actually test
- METR RE-Bench: seven ML research-engineering environments with expert-human attempts; objective optimization and experiment work, typically short feedback loops. Strength: paired human baselines, reproducible artifacts, held-out tasks possible. Missing: ambiguous agendas, long validation, interacting projects and full-scale compute. METR’s original report and paper explicitly contrast 8–32-hour tasks and rapid feedback with months-long lab research; its published suite is only seven environments.
- OpenAI MLE-bench: 75 historical Kaggle challenges, held-out leaderboard scores and bronze-medal thresholds. Tests ML engineering, iteration and resource use across diverse datasets, not selection of important unsolved research questions; leaderboard, scaffolding and data-contamination choices matter.
- OpenAI PaperBench: replication of 20 ICML papers from scratch, author-informed rubrics and partial PhD comparisons. Tests longer integrated engineering, but replicating a known contribution differs from inventing one; published papers/code, rubric gaming and imperfect automated grading constrain inference (paper limitations).
- ResearchGym / conceptual tasks: ResearchGym hides methods from five recent paper tasks while retaining baselines and evaluation; it reports rare standout gains amid unreliable overall performance, but remains tiny and still specifies target metrics. Conceptual-argument ratings probe critique rather than code, but expert agreement and transfer to decisions in a real research programme must be established.
- METR’s September five-task Opus 5.5 assessment: Budget NanoGPT, LMCA, Train a Program, Gaming Bot and open-ended Sunlight, plus questionnaire/interview and a separate inside-lab judgment without disclosed supporting evidence. It is not a purely scoreable-benchmark finding: METR reports incremental improvement even on harder-to-verify LMCA and Sunlight, but no large improvement in foresight, self-generated feedback and taste. It explicitly says the available data cannot distinguish constant from accelerating or decelerating improvement.
Validity traps and the fairest counterargument
METR’s May 2026 report documents saturation of its time-horizon suite and stronger measured performance on less messy tasks; saturation makes some comparisons uninformative, rather than proving underlying research ability stopped improving. Public-task contamination, small task counts, best-of-k runs, attempts to exploit graders, GPU/agent costs, human baseline mismatch and scaffolding changes can move scores without changing autonomous research value. Conversely, bounded engineering is a real input to R&D; reliably automating many such tasks can raise a lab’s pace even if humans continue choosing ideas. Benchmark skepticism does not show the judgment gap is permanent. Chan et al. (2026) propose tracking actual lab time allocation, spending and consequences alongside capabilities, precisely because benchmarks do not identify realized automation.
Discriminating test
Use fresh, blind, non-contaminated and independently judged tasks that require choosing a worthwhile research direction, running and correcting expensive experiments, knowing when to stop, and transferring a validated improvement to a successor model. Publish repeated-run distributions, longer budgets, human/hybrid baselines and all-in experiment/inference costs; then compare realized quality-adjusted lab throughput across generations holding outside inputs constant. Report benchmark ceiling/refresh history, rubric failure and whether improvement survives tasks selected after model training. Neither one score jump nor one judgment failure alone identifies a recursive feedback loop.