TasteVal: a bounded human-beating experiment-design result, not autonomous research agenda setting
TasteVal: a bounded human-beating experiment-design result, not autonomous research agenda setting
Oliver Jaffe and Dane Sherburn, “TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts”, arXiv v2, 6 October 2026 (38 pp). Read 10 October. Connect to prospective-discovery brief, TASTE and RSI economics.
Claim and design. Eight newly designed, withheld AI-R&D tasks include training, data and preference-model problems, with given goal, fast validation and a hidden test. A Researcher chooses the next experiment and interprets returned results; the same fixed Opus 4.8 Coder implements each experiment for model and human Researchers; experiment and report monitors attempt to keep Coder from supplying research judgment or revealing test scores. Each has one H100, serial experiments, 40 H100-busy hours or 120 elapsed hours. Twenty models get six seeds/task; 24 experienced humans (at least two/task) get pre-study literature-review time and the best human per task is baseline. Opus 5.5 refuses one task, and its headline score is on seven, treating refusal as missing; on common seven-task analysis score is 2.25 rather than 2.30. Tasks/details are private to reduce contamination; independent reproducibility is limited. The hidden test scores every experiment, but the researcher sees only validation score; monitors and post hoc model auditor are imperfect safeguards, not independent proof of no leak.
What improved. Opus 5.5’s geometric-mean serial experimental-compute multiplier is 2.30 (bootstrap 95% CI 1.15–4.37) compared with the best-per-task human baseline, meaning it reaches matched test scores with less H100-busy time on this suite, not that it is twice as good at choosing important research questions. Under authors’ cost accounting, $282 per model run versus $9,123 per expert run includes billed model API and GPU at $2.50/hour for both and human time at $119/hour for the expert; this is a task-specific accounting, not a frontier-lab cost estimate. Their fitted compute-multiplier doubling speeds up from ~14 months before December 2025 to ~3 months after (CI 1.7–5), but the final-performance multiplier has no statistically significant break (p=0.12). Models vary scaffolds at early dates; trend fits selected at-release frontier models, and few tasks dominate uncertainty. GPT-5.2 break and extrapolation should not be read as successor generations causing their own faster training.
Strongest objection and best reply. Authors explicitly design for low-noise scores and fast feedback on a single accelerator, omit research-problem selection and group coordination, and find no invented methods among 540 sampled official submissions labeled by another model (composition/modification do occur). Thus this is a direct counterexample to a universal claim that models cannot choose good experiments, not to Dru’s narrower worry about foresight, saturation and unscripted research taste. Better idea selection, experiments in parallel, high-dollar agent budgets and improved scaffolds could make benchmark estimate understate useful automation; the measured efficiency matters even if a human continues setting aims. By contrast the paper’s forecast-model changes (one simulation 51%→88% taste-only singularity) require extending this task-specific slope to real frontier labs, to problem choice, assuming continued trend and little human spending. The paper calls those calculations naive extrapolations; they do not document self-sustaining RSI or loss of control.
Next check. External, prospective seven/eight-task replication with independent review of reports and audit plus harder-to-verify, longer tasks; matched expert–agent–hybrid teams on pre-registered new questions, whole-project gains per all-in researcher+training+experiment dollar and human review hours across at least two successor generations. Compare the distinct endpoint in Si et al.’s executed idea trial and outcome forecasting in Wen et al..