Executed research ideas: apparent novelty did not survive project work
Executed research ideas: apparent novelty did not survive project work
Si, Hashimoto and Yang, “The Ideation–Execution Gap”, original June 2025 study (ICLR 2026). Read 10 October 2026 for research-outcome brief. Forty-three experienced NLP researchers were randomly assigned a source-anonymized human (19) or model-written (24) idea, given three months, spent roughly 100 hours per person on implementation and wrote short papers reviewed blind. Pre-execution expert ratings preferred the model’s ideas for novelty, excitement and likely effectiveness; after execution the AI condition's average rating fell much more on these measures than the human condition. The post-execution difference between human- and model-idea projects is not statistically significant when 43 projects are the independent units. The significant difference is in the pre/post score change; pre/post expectations and final blind paper quality are different endpoints, so do not claim the trial proves human ideas yield better results. The score is still expert review of executed projects, not independent scale-up of verified advances. Expert implementers repaired some details, and topic coverage and number of projects are limited. Contrast TasteVal, which fixes the question and checks hidden test scores for experiment choices, with Wen et al., which predicts outcomes for chosen pairs.