GraphEval: A Lightweight Graph-Based LLM Framework for Idea Evaluation
Tao Feng, Yihang Sun, Jiaxuan You
Abstract
The powerful capabilities of Large Language Models (LLMs) have led to their growing use in evaluating human-generated content, particularly in evaluating research ideas within academic settings. Existing solutions primarily rely on prompt-based LLM methods or fine-tuned lightweight language models for idea evaluation. However, these methods are often unstable and struggle to comprehend the complex semantic information embedded in the ideas, impeding their ability to perform highquality evaluations. To address the above challenges, we propose GraphEval, a lightweight graph-based LLM framework for idea evaluation. Our insight is that a complex idea can be broken down into comprehensible viewpoint-nodes using small prompted LLMs. These viewpoint-nodes can then be linked together through edges created from LLM-based relation extraction and/or BERT similarity scores. The created viewpoint-graph can be used to conveniently propagate scores across viewpoint-nodes to improve the robustness of the idea evaluations. In particular, we propose two lightweight graph-based methods for idea evaluation: (1) GraphEval-LP: a training-free label propagation algorithm that propagates quality labels from known viewpoint-nodes to unknown nodes; (2) GraphEval-GNN: a Graph Neural Network (GNN) that is trained to predict the quality labels given the observed graph with minimal computation resources. Moreover, to overcome LLM's limitation in objectively assessing the novelty of ideas, we further propose a novelty detection model to GraphEval-GNN to enhance its capability in judging idea novelty. Experiments on two datasets show GraphEval improves F1 scores by at least 14% with low computation and API costs. Additionally, GraphEval can effectively detect plagiarized ideas. Our codes for GraphEval is released at https://github.com/ulab-uiuc/GraphEval .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b390862d-c245-4a6f-a51d-9fd65a16e019Cited by top-tier papers6
- The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research IdeasChenglei Si, Tatsunori Hashimoto, Diyi YangICLR 2026 · 60 citations
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- From Automation to Autonomy: A Survey on Large Language Models in Scientific DiscoveryTianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang et al.EMNLP 2025 · 5 citations
- ARCHE: A Novel Task to Evaluate LLMs on Latent Reasoning Chain ExtractionPengze Li, Jiaqi Liu, Junchi Yu, Lihao Liu et al.AAAI 2026 · 1 citation
- Beyond Noise: Characterizing Creative Potential in Unverifiable LLM HallucinationsYu Yan, Chunhong Zhang, Haiyu Zhao, Ziyang Zeng et al.ACL 2026
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
Related papers
- Can Large Language Models Unlock Novel Scientific Research Ideas?Sandeep Kumar, Tirthankar Ghosal, Vinayak Goyal, Asif EkbalEMNLP 2025 · 3 citations
- Semantic-Eval : A Semantic Comprehension Evaluation Framework for Large Language Models Generation without TrainingShusheng Li, Jiale Li, Yifei Qu, Xinwei Shi et al.ACL 2025
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP ResearchersChenglei Si, Diyi Yang, Tatsunori HashimotoICLR 2025
- InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning ProblemShuofei Qiao, Yunxiang Wei, Xuehai Wang, Bin Wu et al.ICML 2026
- GraphMind: Unveiling Scientific Reasoning through Contextual Graphs for Novelty AssessmentItalo Luis da Silva, Hanqi Yan, Lin Gui, Yulan HeKDD 2026 · 1 citation
