Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
Gal Yona, Roee Aharoni, Mor Geva
Abstract
Factual questions can typically be answered correctly at different levels of granularity. For example, both " August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?". Standard question answering (QA) evaluation protocols, however, do not take this into account explicitly and instead compare a predicted answer against reference answers of a single granularity level. In this work, we propose GRANOLA QA, a novel evaluation setting where a predicted answer is evaluated in terms of accuracy and informativeness against a set of multi-granularity answers. We present a simple methodology for enriching existing datasets with multi-granularity answers, and create GRANOLA-EQ, a multigranularity version of the ENTITYQUESTIONS dataset. 1 We evaluate models using a range of decoding methods on GRANOLA-EQ, including a new algorithm called Decoding with Response Aggregation (DRAG), that is geared towards aligning the answer granularity with the model's uncertainty. Our experiments show that large language models with standard decoding methods tend to generate specific answers, which are often incorrect. In contrast, when evaluated on multi-granularity answers, DRAG yields a nearly 20 point increase in accuracy on average, which further increases for rare entities, revealing that standard evaluation and decoding schemes may underestimate the knowledge encapsulated in language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97bf6ad6-e145-45cc-8d43-167693a23e85Cited by top-tier papers7
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal et al.EMNLP 2024 · 53 citations
- Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric FactualityNitay Calderon, Eyal Ben-David, Zorik Gekhman, Eran Ofek et al.ICML 2026 · 8 citations
- Abstraction Alignment: Comparing Model-Learned and Human-Encoded Conceptual RelationshipsAngie W. Boggust, Hyemin Bang, Hendrik Strobelt, Arvind SatyanarayanCHI 2025 · 4 citations
- Estimating Knowledge in Large Language Models Without Generating a Single TokenDaniela Gottesman, Mor GevaEMNLP 2024 · 2 citations
- Evaluating LLM Uncertainty in Long-Form Generation Using Deterministic Ground TruthIdo Amit, Ido Galil, Ran El-YanivICML 2026
Builds on9
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu et al.ICML 2024 · 406 citations
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das et al.ACL 2023 · 233 citations
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 162 citations
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 121 citations
- Evaluating Open-Domain Question Answering in the Era of Large Language ModelsEhsan Kamalloo, Nouha Dziri, Charles L. A. Clarke, Davood RafieiACL 2023 · 96 citations
Related papers
- An Empirical Study of Evaluating Long-form Question AnsweringNing Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke et al.SIGIR 2025 · 2 citations
- Integrative Decoding: Improving Factuality via Implicit Self-consistencyYi Cheng, Xiao Liang, Yeyun Gong, Wen Xiao et al.ICLR 2025 · 1 citation
- CIRAG: Retrieval-Augmented Language Model with Collective IntelligenceChenxu Cui, Haihui Fan, Jinchao Zhang, Lin Shen et al.SIGIR 2025 · 5 citations
- Multi-granularity Temporal Question Answering over Knowledge GraphsZiyang Chen, Jinzhi Liao, Xiang ZhaoACL 2023 · 36 citations
- OBELLA: Open the Book for Evaluating Long-Form Large Language Model Answers in Open-Domain Question AnsweringTianyu Ren, Zhaoyu Zhang, Hui Wang, Karen RaffertySIGIR 2025
