GRAB: A Challenging Graph Analysis Benchmark for Large Multimodal Models
Jonathan Roberts, Kai Han, Samuel Albanie
Abstract
Large multimodal models (LMMs) have exhibited proficiencies across many visual tasks. Although numerous well-known benchmarks exist to evaluate model performance, they increasingly have insufficient headroom. As such, there is a pressing need for a new generation of benchmarks challenging enough for the next generation of LMMs. One area that LMMs show potential is graph analysis, specifically, the tasks an analyst might typically perform when interpreting figures such as estimating the mean, intercepts or correlations of functions and data series. In this work, we introduce GRAB, a graph analysis benchmark, fit for current and future frontier LMMs. Our benchmark is predominantly synthetic, ensuring high-quality, noise-free questions. GRAB is comprised of 3284 questions, covering five tasks and 23 graph properties. We evaluate 20 LMMs on GRAB, finding it to be a challenging benchmark, with the highest performing model attaining a score of just 21.0%. Finally, we conduct various ablations to investigate where the models succeed and struggle. We release GRAB and a lightweight GRAB-Lite to encourage progress in this important, growing domain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ba84b6a-85da-4491-b44b-8d0c7e3d1b5dCited by top-tier papers2
- ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal ModelsJonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma, Akash Gupta et al.ICML 2026
- Needle Threading: Can LLMs Follow Threads Through Near-Million-Scale Haystacks?Jonathan Roberts, Kai Han, Samuel AlbanieICLR 2025
Builds on8
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang et al.NeurIPS 2024 · 1,029 citations
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong et al.NeurIPS 2024 · 858 citations
Related papers
- Evaluating LLMs on Large-Scale Graph Property Estimation via Random WalksSunil Kumar Maurya, Xin LiuACL 2026
- Text2Analysis: A Benchmark of Table Question Answering with Advanced Data Analysis and Unclear QueriesXinyi He, Mengyu Zhou, Xinrun Xu, Xiaojun Ma et al.AAAI 2024 · 48 citations
- Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code GenerationJiawei Zhou, Chi Zhang, Xiang Feng, Qiming Zhang et al.ACL 2026 · 2 citations
- MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language ModelsPeng Xia, Siwei Han, Shi Qiu, Yiyang Zhou et al.ICLR 2025
- VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal ModelsYuntao Du, Yiming Wang, Renshuo Yuan, Jincheng Yue et al.CVPR 2026
