TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
Two researchers built eight withheld AI research tasks with clear goals and fast experimental feedback. Experienced humans and AI models both chose what experiment to run next; the same fixed coding assistant implemented their choices. Neither side saw the hidden test scores used for evaluation. On the seven tasks its top model attempted, the model reached comparable hidden-test quality using 2.30 times less serial experiment compute than the best human assigned to each task. The comparison favors the human baseline by choosing its best performer per task, though the model did not complete the eighth task.
The authors argue that experiment choice, not just coding, is improving quickly. Their estimated efficiency trend steepens markedly in recent models, while their separate measure of final research quality shows no statistically significant change in trend. Both are fits to a small, purpose-built task suite, not observations that a frontier lab’s total research output is accelerating across model generations.
The tasks provide the question, success metric and cheap feedback. They do not test who picks a worthwhile unsolved problem, organizes a lab or validates a discovery after months of costly work. An audit classified no sampled submissions as invented methods, though models combined and modified existing ones. A separate executed-idea study found that AI ideas’ high initial ratings weakened after implementation. The next decisive comparison would follow independently chosen agent, human and hybrid projects to verified results, counting all experiment compute and human review across successive models.