Evaluating LLMs on Large-Scale Graph Property Estimation via Random Walks
Sunil Kumar Maurya, Xin Liu
Abstract
With the rapidly improving reasoning abilities of Large Language Models (LLMs), there is also a rising demand to use them in a wide variety of domains. This brings about the need to carefully evaluate the limits of the capabilities of these models with various tests and benchmarks. Graph structures are ubiquitous in real-world data, and are often used to represent and analyze relationship patterns within data. Many benchmarks have already been proposed in the graph literature to test the reasoning ability of LLMs to follow and execute graph algorithms. However, due to the limited context length of LLMs, these benchmarks consist of very small graphs. In real-world data, the size of graphs can be significantly larger, and in many cases, not fully accessible. In this paper, we examine a class of problems that arises with very large graphs having limited accessibility. We propose a large graph benchmark dataset, EstGraph, and introduce four distinct tasks designed to estimate large graph properties. We evaluate the reasoning abilities of LLMs on these tasks using a wide variety of graph datasets. In addition, we provide taskspecific prompt constructions based on random walk sampling of large graphs (up to millions of nodes) that effectively convey sufficient information to LLMs within the limits of context length. Source code and datasets are available at https://zenodo.org/records/19632942
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7ab4df7-52f3-4320-b50a-49a864ab3e7aBuilds on5
- Can Language Models Solve Graph Problems in Natural Language?Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan et al.NeurIPS 2023 · 420 citations
- Talk like a Graph: Encoding Graphs for Large Language ModelsBahare Fatemi, Jonathan Halcrow, Bryan PerozziICLR 2024 · 194 citations
- Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph InferenceBen Finkelshtein, Silviu Cucerzan, Sujay Kumar Jauhar, Ryen W WhiteICLR 2026 · 5 citations
- How Do Large Language Models Understand Graph Patterns? A Benchmark for Graph Pattern ComprehensionXinnan Dai, Haohao Qu, Yifei Shen, Bohang Zhang et al.ICLR 2025 · 1 citation
- GraphArena: Evaluating and Exploring Large Language Models on Graph ComputationJianheng Tang, Qifan Zhang, Yuhan Li, Nuo Chen et al.ICLR 2025
Related papers
- Can Graph Descriptive Order Affect Solving Graph Problems with LLMs?Yuyao Ge, Shenghua Liu, Baolong Bi, Yiwei Wang et al.ACL 2025
- MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular GraphsChristoph Bartmann, Johannes Schimunek, Mykyta Ielanskyi, Philipp Seidl et al.ICLR 2026 · 5 citations
- <tt>G1</tt>: Teaching LLMs to Reason on Graphs with Reinforcement LearningXiaojun Guo, Ang Li, Yifei Wang, Stefanie Jegelka et al.NeurIPS 2025 · 16 citations
- LLM4DyG: Can Large Language Models Solve Spatial-Temporal Problems on Dynamic Graphs?Zeyang Zhang, Xin Wang, Ziwei Zhang, Haoyang Li et al.KDD 2024 · 32 citations
- GRAB: A Challenging Graph Analysis Benchmark for Large Multimodal ModelsJonathan Roberts, Kai Han, Samuel AlbanieICCV 2025
