Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
Gal Yona, Roee Aharoni, Mor Geva
摘要
Factual questions can typically be answered correctly at different levels of granularity. For example, both " August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?". Standard question answering (QA) evaluation protocols, however, do not take this into account explicitly and instead compare a predicted answer against reference answers of a single granularity level. In this work, we propose GRANOLA QA, a novel evaluation setting where a predicted answer is evaluated in terms of accuracy and informativeness against a set of multi-granularity answers. We present a simple methodology for enriching existing datasets with multi-granularity answers, and create GRANOLA-EQ, a multigranularity version of the ENTITYQUESTIONS dataset. 1 We evaluate models using a range of decoding methods on GRANOLA-EQ, including a new algorithm called Decoding with Response Aggregation (DRAG), that is geared towards aligning the answer granularity with the model's uncertainty. Our experiments show that large language models with standard decoding methods tend to generate specific answers, which are often incorrect. In contrast, when evaluated on multi-granularity answers, DRAG yields a nearly 20 point increase in accuracy on average, which further increases for rare entities, revealing that standard evaluation and decoding schemes may underestimate the knowledge encapsulated in language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal 等EMNLP 2024 · 被引用 53 次
- Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric FactualityNitay Calderon, Eyal Ben-David, Zorik Gekhman, Eran Ofek 等ICML 2026 · 被引用 8 次
- Abstraction Alignment: Comparing Model-Learned and Human-Encoded Conceptual RelationshipsAngie W. Boggust, Hyemin Bang, Hendrik Strobelt, Arvind SatyanarayanCHI 2025 · 被引用 4 次
- Estimating Knowledge in Large Language Models Without Generating a Single TokenDaniela Gottesman, Mor GevaEMNLP 2024 · 被引用 2 次
- Evaluating LLM Uncertainty in Long-Form Generation Using Deterministic Ground TruthIdo Amit, Ido Galil, Ran El-YanivICML 2026
它引用的顶会 Paper9
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu 等ICML 2024 · 被引用 406 次
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das 等ACL 2023 · 被引用 233 次
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 被引用 162 次
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 被引用 121 次
- Evaluating Open-Domain Question Answering in the Era of Large Language ModelsEhsan Kamalloo, Nouha Dziri, Charles L. A. Clarke, Davood RafieiACL 2023 · 被引用 96 次
相关 Paper
- An Empirical Study of Evaluating Long-form Question AnsweringNing Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke 等SIGIR 2025 · 被引用 2 次
- Integrative Decoding: Improving Factuality via Implicit Self-consistencyYi Cheng, Xiao Liang, Yeyun Gong, Wen Xiao 等ICLR 2025 · 被引用 1 次
- CIRAG: Retrieval-Augmented Language Model with Collective IntelligenceChenxu Cui, Haihui Fan, Jinchao Zhang, Lin Shen 等SIGIR 2025 · 被引用 5 次
- Multi-granularity Temporal Question Answering over Knowledge GraphsZiyang Chen, Jinzhi Liao, Xiang ZhaoACL 2023 · 被引用 36 次
- OBELLA: Open the Book for Evaluating Long-Form Large Language Model Answers in Open-Domain Question AnsweringTianyu Ren, Zhaoyu Zhang, Hui Wang, Karen RaffertySIGIR 2025
