Evaluating Open-Domain Question Answering in the Era of Large Language Models
Ehsan Kamalloo, Nouha Dziri, Charles L. A. Clarke, Davood Rafiei
Abstract
Lexical matching remains the de facto evaluation method for open-domain question answering (QA). Unfortunately, lexical matching fails completely when a plausible candidate answer does not appear in the list of gold answers, which is increasingly the case as we shift from extractive to generative models. The recent success of large language models (LLMs) for QA aggravates lexical matching failures since candidate answers become longer, thereby making matching with the gold answers even more challenging. Without accurate evaluation, the true progress in open-domain QA remains unknown. In this paper, we conduct a thorough analysis of various open-domain QA models, including LLMs, by manually evaluating their answers on a subset of NQ-OPEN, a popular benchmark. Our assessments reveal that while the true performance of all models is significantly underestimated, the performance of the InstructGPT (zero-shot) LLM increases by nearly +60%, making it on par with existing top models, and the InstructGPT (few-shot) model actually achieves a new state-of-the-art on NQ-OPEN. We also find that more than 50% of lexical matching failures are attributed to semantically equivalent answers. We further demonstrate that regex matching ranks QA models consistent with human judgments, although still suffering from unnecessary strictness. Finally, we demonstrate that automated evaluation models are a reasonable surrogate for lexical matching in some circumstances, but not for long-form answers generated by LLMs. The automated models struggle in detecting hallucinations in LLM answers and are thus unable to evaluate LLMs. At this time, there appears to be no substitute for human evaluation. 1 What was the city of Beijing previously known as?
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8058d8c1-b640-4d65-a2c5-423fc61556caCited by top-tier papers67
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SuRe: Summarizing Retrievals using Answer Candidates for Open-domain QA of LLMsJaehyung Kim, Jaehyun Nam, Sangwoo Mo, Jongjin Park et al.ICLR 2024 · 89 citations
- Improving Automatic VQA Evaluation Using Large Language ModelsOscar Mañas, Benno Krojer, Aishwarya AgrawalAAAI 2024 · 59 citations
- ArkVale: Efficient Generative LLM Inference with Recallable Key-Value EvictionRenze Chen, Zhuofeng Wang, Beiquan Cao, Tong Wu et al.NeurIPS 2024 · 56 citations
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal et al.EMNLP 2024 · 53 citations
Builds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher et al.ICLR 2020 · 322 citations
- End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question AnsweringDevendra Singh Sachan, Siva Reddy, William L. Hamilton, Chris Dyer et al.NeurIPS 2021 · 197 citations
Related papers
- An Empirical Study of Evaluating Long-form Question AnsweringNing Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke et al.SIGIR 2025 · 2 citations
- Factuality of Large Language Models: A SurveyYuxia Wang, Minghan Wang, Muhammad Arslan Manzoor, Fei Liu et al.EMNLP 2024 · 23 citations
- OBELLA: Open the Book for Evaluating Long-Form Large Language Model Answers in Open-Domain Question AnsweringTianyu Ren, Zhaoyu Zhang, Hui Wang, Karen RaffertySIGIR 2025
- The Illusion of Progress: Re-evaluating Hallucination Detection in LLMsDenis Janiak, Jakub Binkowski, Albert Sawczyn, Bogdan Gabrys et al.EMNLP 2025
- Enabling Large Language Models to Generate Text with CitationsTianyu Gao, Howard Yen, Jiatong Yu, Danqi ChenEMNLP 2023 · 152 citations
