CBench: Towards Better Evaluation of Question Answering Over Knowledge Graphs
Abdelghny Orogat, Isabelle Liu, Ahmed El-Roby
摘要
Recently, there has been an increase in the number of knowledge graphs that can be only queried by experts. However, describing questions using structured queries is not straightforward for non-expert users who need to have sufficient knowledge about both the vocabulary and the structure of the queried knowledge graph, as well as the syntax of the structured query language used to describe the user's information needs. The most popular approach introduced to overcome the aforementioned challenges is to use natural language to query these knowledge graphs. Although several question answering benchmarks can be used to evaluate question-answering systems over a number of popular knowledge graphs, choosing a benchmark to accurately assess the quality of a question answering system is a challenging task. In this paper, we introduce CBench, an extensible, and more informative benchmarking suite for analyzing benchmarks and evaluating question answering systems. CBench can be used to analyze existing benchmarks with respect to several fine-grained linguistic, syntactic, and structural properties of the questions and queries in the benchmark. We show that existing benchmarks vary significantly with respect to these properties deeming choosing a small subset of them unreliable in evaluating QA systems. Until further research improves the quality and comprehensiveness of benchmarks, CBench can be used to facilitate this evaluation using a set of popular benchmarks that can be augmented with other user-provided benchmarks. CBench not only evaluates a question answering system based on popular single-number metrics but also gives a detailed analysis of the linguistic, syntactic, and structural properties of answered and unanswered questions to better help the developers of question answering systems to better understand where their system excels and where it struggles.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Maestro: Automatic Generation of Comprehensive Benchmarks for Question Answering Over Knowledge GraphsAbdelghny Orogat, Ahmed El-RobySIGMOD 2023 · 被引用 10 次
- Is Complex Query Answering Really Complex?Cosimo Gregucci, Bo Xiong, Daniel Hernández, Lorenzo Loconte 等ICML 2025
- Can LLMs Evaluate Complex Attribution in QA? Automatic Benchmarking using Knowledge GraphsNan Hu, Jiaoyan Chen, Yike Wu, Guilin Qi 等ACL 2025
- TabReX: Tabular Referenceless eXplainable EvaluationTejas Anvekar, Junha Park, Aparna Garimella, Vivek GuptaACL 2026
- Generative Language Models for Paragraph-Level Question GenerationAsahi Ushio, Fernando Alva-Manchego, José Camacho-ColladosEMNLP 2022 · 被引用 30 次
