Forecasting experiments researchers already chose is not choosing the research agenda
Forecasting experiments researchers already chose is not choosing the research agenda
Wen et al., “Predicting Empirical AI Research Outcomes with Language Models”, June 2025 original preprint; checked against full PDF 10 October 2026. Predicts which of two already described methods works better on at least three published-paper benchmarks, aggregating winners over metrics. The authors trained a GPT-4.1-based system on 6,000 historical idea pairs and tested on 1,585 human-verified pairs that include at least one idea published after the base model’s June 2024 cutoff. Accuracy is 77% on this large test set, not a human-versus-model comparison. Against 25 NLP experts on a 45-pair subset (five experts per prediction), the fine-tuned model plus paper retrieval scored 64.4% versus 48.9% for expert aggregation. Off-the-shelf frontier models with retrieval performed near chance, so the fine-tuned system result is not a generic out-of-the-box model result. On 35 unpublished, executed research ideas from another project, its accuracy was 63.6%, without matched experts; the sample reduces public-contamination concerns, but is not a live preregistered prediction-and-selection trial.
A real challenge to saying that outcome judgment always requires humans, and to relying solely on subjective idea ratings. Conversely, the pair and benchmarks are given, output is a binary outcome drawn from paper experiments, and training and filtering matter. Reuse of benchmark conventions, publication selection and 45 expert-comparison items restrict generalization to choosing important new problems or improving successor models. Compare TasteVal on hidden-test iterative experiment choice and Si et al. on executing assigned ideas. Question hub: research taste and discovery.