Maestro: Automatic Generation of Comprehensive Benchmarks for Question Answering Over Knowledge Graphs
Abdelghny Orogat, Ahmed El-Roby
Abstract
Recently, there has been an upsurge in the number of knowledge graphs (KG) that can only be accessed by experts. Non-expert users lack an adequate understanding of the queried knowledge graph's vocabulary and structure, as well as the syntax of the structured query language used to express the user's information needs. To increase the user base of these KGs, a set of Question Answering (QA) systems that use natural language to query these knowledge graphs have been introduced. However, finding a benchmark that accurately evaluates the quality of a QA system is a difficult task due to (1) the high degree of variation in the fine-grained properties among the existing benchmarks, (2) the static nature of the existing benchmarks versus the evolving nature of KGs, and (3) the limited number of KGs targeted by existing benchmarks, which hinders the usability of QA systems in real-world deployment over KGs that are different from those that were used in the evaluation of the QA systems. In this paper, we introduce Maestro, a benchmark generation system for question answering over knowledge graphs. Maestro can generate a new benchmark for any KG given the KG and, optionally, a text corpus that covers this KG. The benchmark generated by Maestro is guaranteed to cover all the properties of the natural language questions and queries that were encountered in the literature as long as the targeted KG includes these properties. Maestro also generates high-quality natural language questions with various utterances that are on par with manually-generated ones to better evaluate QA systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ad8b1b3a-1d36-4b4b-b729-2a21ab4270aeCited by top-tier papers2
- Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMsReham Omar, Omij Mangukiya, Essam MansourSIGMOD 2025 · 9 citations
- LLMLog: Advanced Log Template Generation via LLM-driven Multi-Round AnnotationFei Teng, Haoyang Li, Lei ChenVLDB 2025 · 2 citations
Related papers
- CBench: Towards Better Evaluation of Question Answering Over Knowledge GraphsAbdelghny Orogat, Isabelle Liu, Ahmed El-RobyVLDB 2021 · 17 citations
- Can LLMs Evaluate Complex Attribution in QA? Automatic Benchmarking using Knowledge GraphsNan Hu, Jiaoyan Chen, Yike Wu, Guilin Qi et al.ACL 2025
- A Universal Question-Answering Platform for Knowledge GraphsReham Omar, Ishika Dhall, Panos Kalnis, Essam MansourSIGMOD 2023 · 48 citations
- Is Complex Query Answering Really Complex?Cosimo Gregucci, Bo Xiong, Daniel Hernández, Lorenzo Loconte et al.ICML 2025
- Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and OpportunitiesChuangtao Ma, Yongrui Chen, Tianxing Wu, Arijit Khan et al.EMNLP 2025 · 6 citations
