GQLBench: A Large-Scale Cross-Domain, Cross-Dialect Benchmark for NL2GQL
Yanning Su, Yuhang Zhou, Yang Fang, Sen Liu, Guangnan Ye, Hongfeng Chai
Abstract
. Abstract Despite growing interest in NL2GQL, benchmarking progress has been constrained by the lack of resources that are simultaneously large-scale, cross-domain, and cross-dialect. To address this gap, we present GQLBench , a new benchmark built through an automated and scalable framework that integrates NL2SQL-to-NL2GQL conversion with graph-native data generation. GQLBench supports execution-based evaluation on both Cypher and ISO GQL, covering hundreds of graph databases and over 20k natural language questions for each dialect. By combining converted data from mature NL2SQL resources with synthetic graph-specific queries, it captures both schema diversity from real-world relational sources and graph-native reasoning challenges, including long paths and cycles. Beyond overall performance comparison, GQLBench also enables fine-grained evaluation across dialects, graph patterns, and query complexity. Experiments on advanced LLMs show that even strong proprietary models struggle on GQLBench, with gemini-3-flash achieving only 35.40% average execution accuracy across the two di-alects. Our data and code are available at https://github.com/qxssadf/GQLBench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language ModelsBernal Jimenez Gutierrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga et al.NeurIPS 2024 · 395 citations
- GraphQ IR: Unifying the Semantic Parsing of Graph Query Languages with One Intermediate RepresentationLunyiu Nie, Shulin Cao, Jiaxin Shi, Jiuding Sun et al.EMNLP 2022 · 20 citations
- CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM EraYanlin Feng, Simone Papicchio, Sajjadur RahmanACL 2025
- From RAG to Memory: Non-Parametric Continual Learning for Large Language ModelsBernal Jiménez Gutiérrez, Yiheng Shu, Weijian Qi, Sizhe Zhou et al.ICML 2025
Related papers
- NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL SolutionsShizheng Hou, Wenqi Pei, Nuo Chen, Quang-Trung Ta et al.VLDB 2026 · 1 citation
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon et al.ICLR 2026 · 13 citations
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal ModelsJingxuan Wei, Caijun Jia, Xi Bai, Xinglong Xu et al.CVPR 2026 · 7 citations
- GraphOmni: A Comprehensive and Extensible Benchmark Framework for Large Language Models on Graph-theoretic TasksHao Xu, Xiangru Jian, Xinjian Zhao, Wei Pang et al.ICLR 2026 · 6 citations
- LogiConBench: Benchmarking Logical Consistencies of LLMsZheng Chen, Chuan Zhou, Fengxiang Cheng, Tin Po Yip et al.ICLR 2026
