Does AI research taste predict validated discoveries, rather than expert approval?
Does AI research taste predict validated discoveries, rather than expert approval?
Brief, completed 10 October 2026. Question: where has machine selection of new research directions been linked to verified downstream results at matched resources and human baselines? A good answer distinguishes expert preference, outcome forecasting for a given idea, choosing an experiment for a given research goal, executing an assigned new project, and across-generation capability compounding. Earlier TASTE measures selected expert preferences, not success; feedback synthesis measures bounded optimization and supervised task delegation; Dru’s question concerns saturation, foresight and taste.
Search and short primary-source list
- Chosen experiments → new hidden outcomes: search for 2026 prospective AI-chosen experiments / expert human evaluation. Jaffe & Sherburn, TasteVal, 6 October 2026, preprint; eight withheld research-engineering tasks and 24 experts, seven-task headline model, fixed coder, hidden test.
- Chosen research idea → executed result: search for randomized expert execution of AI and human proposals. Si, Hashimoto & Yang’s randomized NLP idea-execution study, 43 researchers, 19 human/24 machine-sourced projects. Unlike a scoreable experiment-selection benchmark, people run the ideas to projects; post-execution between-arm difference is not significant.
- Predicted outcome → project selection: search for prospective outcome prediction and models’ research judgment. Wen et al. test 1,585 post-cutoff verified pairs, with expert comparison on just 45; human challengers here forecast fixed pair outcomes, not choose and execute new programmes.
- Feedback and gaming → compounding: rechecked original TasteVal full text, October 4 synthesis and original Cunningham et al. economic threshold model. TasteVal’s automatically checked test scores reduce a pure subjective-rating weakness, but monitors and model audits are not proof of leakage absence. No matched successor-generation lab series of human decisions, verified new research outcomes and all-in input costs in this source shortlist. A finite search cannot rule out another study.
Result and named disagreement
There is now stronger evidence against a universal taste bottleneck: in TasteVal, the tested frontier model reaches given hidden-test target quality with 2.30× less serial experiment compute than the best human per task on seven scoped tasks; the estimated compute-efficiency trend steepens. Wen et al. show a specialized fine-tuned forecaster can beat a small expert sample on already selected paper-method pairs. Both are machine judgment rather than just code generation. This pushes against Dru’s reading if it is that a model cannot reliably select the next experiment and supports the authors’ claim that a lab with many measurable subproblems could see material gains with humans still picking objectives.
It does not settle his narrower objection: TasteVal problems, target metrics, low-noise feedback, short experiment budgets and single-GPU setup are fixed by humans; its authors find no invented methods in a model-classified submission sample, its top model refuses one task and its final-performance trend shows no significant breakpoint. Si et al. found the pre-execution ratings of model-proposed NLP ideas fell more sharply after execution, with neither side significantly better after execution. The results are compatible, not a contradiction: one measures how well an agent searches within a known target, another how a generated agenda survives expensive execution. A numerical singularity forecast in TasteVal additionally requires extrapolation from serial compute efficiency to real R&D productivity and to successive generations at constant outside inputs; no observation here establishes that chain.
What changes confidence next: blind external test with preselected new, consequential directions; independent teams randomly assigned an agent, expert or hybrid researcher, matched candidate pools and token/human/GPU/training budgets; prespecified out-of-sample consequences at 3–12 months and adverse outcomes; then compare verified per-all-in-dollar advances and review time across two successor models. Also compare the effect of widening the researcher’s mandate from optimizing an agreed metric to selecting one. This distinguishes a genuine expert-taste crossover from benchmark optimization without dismissing the latter.
Feed choice Saturday: post the 38-page TasteVal paper alone, summary sufficient without opening, because it is a new primary head-to-head and the best disagreement with the one topic Dru discussed. Do not post 2025 idea-execution study on the same morning or revive the already-dismissed lab delegation piece. Backlog: Si trial as a weekend counterpoint only if TasteVal is opened/discussed/hearted; Wen as a short weekday original-paper contrast only on a follow-up about forecasting vs discovery. If TasteVal also sits unopened, keep research in notes and leave feed quiet next night.