Forecasting experiments researchers already chose is not choosing the research agenda

#topic #rsi #original-evaluation #primary-source

Forecasting experiments researchers already chose is not choosing the research agenda

Wen et al., “Predicting Empirical AI Research Outcomes with Language Models”, June 2025 original preprint; checked against full PDF 10 October 2026. Predicts which of two already described methods works better on at least three published-paper benchmarks, aggregating winners over metrics. The authors trained a GPT-4.1-based system on 6,000 historical idea pairs and tested on 1,585 human-verified pairs that include at least one idea published after the base model’s June 2024 cutoff. Accuracy is 77% on this large test set, not a human-versus-model comparison. Against 25 NLP experts on a 45-pair subset (five experts per prediction), the fine-tuned model plus paper retrieval scored 64.4% versus 48.9% for expert aggregation. Off-the-shelf frontier models with retrieval performed near chance, so the fine-tuned system result is not a generic out-of-the-box model result. On 35 unpublished, executed research ideas from another project, its accuracy was 63.6%, without matched experts; the sample reduces public-contamination concerns, but is not a live preregistered prediction-and-selection trial.

A real challenge to saying that outcome judgment always requires humans, and to relying solely on subjective idea ratings. Conversely, the pair and benchmarks are given, output is a binary outcome drawn from paper experiments, and training and filtering matter. Reuse of benchmark conventions, publication selection and 45 expert-comparison items restrict generalization to choosing important new problems or improving successor models. Compare TasteVal on hidden-test iterative experiment choice and Si et al. on executing assigned ideas. Question hub: research taste and discovery.