GraphOmni: A Comprehensive and Extensible Benchmark Framework for Large Language Models on Graph-theoretic Tasks
Hao Xu, Xiangru Jian, Xinjian Zhao, Wei Pang, Chao Zhang, Suyuchen Wang, Qixin Zhang, Zhengyuan Dong, Joao Monteiro, Bang Liu, Qiuzhuang Sun, Tianshu Yu
Abstract
This paper introduces GraphOmni, a comprehensive benchmark designed to evaluate the reasoning capabilities of LLMs on graph-theoretic tasks articulated in natural language. GraphOmni spans diverse graph types, serialization formats, and prompting schemes, substantially extending upon prior efforts in both scope and depth. Through systematic evaluation, we uncover critical interactions among these dimensions, revealing their decisive impact on model performance. Our experiments show that state-of-the-art closed-source models such as Claude-3.5 and o4-mini consistently lead overall, yet still leave considerable headroom, while open-source models display pronounced sensitivity to various design choices. Beyond the standard scope, larger graphs, real-world graphs, and additional NP-hard tasks are further discussed. We further analyze efficiency via output token usage, highlighting cost–accuracy trade-offs, and introduce a reinforcement learning-based optimizer that adaptively selects factor combinations, reducing evaluation cost by 75% while retaining strong accuracy. This flexible and extensible benchmark not only deepens understanding of LLM performance on structured graph reasoning but also establishes a robust foundation for advancing model design and evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af4f2d3b-a36b-48d8-b6a9-887747a375a7Cited by top-tier papers2
- <tt>G1</tt>: Teaching LLMs to Reason on Graphs with Reinforcement LearningXiaojun Guo, Ang Li, Yifei Wang, Stefanie Jegelka et al.NeurIPS 2025 · 16 citations
- The Underappreciated Power of Vision Models for Graph Structural UnderstandingXinjian Zhao, Wei Pang, Zhongkai Xue, Xiangru Jian et al.NeurIPS 2025 · 7 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
Related papers
- GraphCogent: Mitigating LLMs' Working Memory Constraints via Multi-Agent Collaboration in Complex Graph UnderstandingRongzheng Wang, Shuang Liang, Qizhi Chen, Yihong Huang et al.WWW 2026 · 4 citations
- GQLBench: A Large-Scale Cross-Domain, Cross-Dialect Benchmark for NL2GQLYanning Su, Yuhang Zhou, Yang Fang, Sen Liu et al.ACL 2026
- How Do Large Language Models Understand Graph Patterns? A Benchmark for Graph Pattern ComprehensionXinnan Dai, Haohao Qu, Yifei Shen, Bohang Zhang et al.ICLR 2025 · 1 citation
- Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph InferenceBen Finkelshtein, Silviu Cucerzan, Sujay Kumar Jauhar, Ryen W WhiteICLR 2026 · 5 citations
- Evaluating LLMs on Large-Scale Graph Property Estimation via Random WalksSunil Kumar Maurya, Xin LiuACL 2026
