The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
Chenglei Si, Tatsunori Hashimoto, Diyi Yang
摘要
Large Language Models (LLMs) have shown promise in accelerating the scientific research pipeline. A key capability for this process is the ability to generate novel research ideas, and prior studies have found settings in which LLM-generated research ideas were judged as more novel than human-expert ideas. However, a good idea should not simply appear to be novel, it should also result in better research after being executed. To test whether AI-generated ideas lead to better research outcomes, we conduct an execution study by recruiting 43 expert researchers to execute randomly-assigned ideas, either written by experts or generated by an LLM. Each expert spent over 100 hours implementing the idea and wrote a 4-page short paper to document the experiments. All the executed projects are then reviewed blindly by expert NLP researchers. Comparing the review scores of the same ideas before and after execution, the scores of the LLM-generated ideas decrease significantly more than expert-written ideas on all evaluation metrics (novelty, excitement, effectiveness, and overall; p < 0.05), closing the gap between LLM and human ideas observed at the ideation stage. When comparing the aggregated review scores from the execution study, we even observe that for many metrics there is a flip in rankings where human ideas score higher than LLM ideas. This ideation-execution gap highlights the limitations of current LLMs in generating truly effective research ideas and the challenge of evaluating research ideas in the absence of execution outcomes. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Towards Execution-Grounded Automated AI ResearchChenglei Si, Zitong Yang, Yejin Choi, Emmanuel J Candes 等ICML 2026 · 被引用 12 次
- Towards AI as Colleagues: Multi-Agent System Improves Structured Ideation ProcessesKexin Quan, Dina Albassam, Mengke Wu, Zijian Ding 等CHI 2026 · 被引用 2 次
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
相关 Paper
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP ResearchersChenglei Si, Diyi Yang, Tatsunori HashimotoICLR 2025
- Can Large Language Models Unlock Novel Scientific Research Ideas?Sandeep Kumar, Tirthankar Ghosal, Vinayak Goyal, Asif EkbalEMNLP 2025 · 被引用 3 次
- Predicting Empirical AI Research Outcomes with Language ModelsJiaxin Wen, Chenglei Si, Yueh-Han Chen, He He 等NeurIPS 2025 · 被引用 18 次
- Timing Matters: How Using LLMs at Different Timings Influences Writers' Perceptions and Ideation Outcomes in AI-Assisted IdeationPeinuan Qin, Chi-Lan Yang, Jingshu Li, Jing Wen 等CHI 2025 · 被引用 19 次
- AAAR-1.0: Assessing AI's Potential to Assist ResearchRenze Lou, Hanzi Xu, Sijia Wang, Jiangshu Du 等ICML 2025
