OBELLA: Open the Book for Evaluating Long-Form Large Language Model Answers in Open-Domain Question Answering
Tianyu Ren, Zhaoyu Zhang, Hui Wang, Karen Rafferty
Abstract
Reliable factuality evaluation is critical for the iterative development of open-domain question answering (ODQA) systems, especially given the rise of large language models (LLMs) and their propensity for hallucination. However, state-of-the-art (SOTA) automatic metrics, which are mostly supervised, remain notably less reliable than humans. In this paper, we find two key challenges behind this gap: (1) length distribution mismatch between lengthy LLM answers and shorter training answers used by current metrics; and (2) reference incompleteness, where current metrics often misjudge valid system answers absent from given references-a challenge worsened by the diversity of LLM outputs. To address these issues, we present a new ODQA factuality evaluation dataset called OBELLA (Open-Book Evaluation for Long-form LLM Answers). OBELLA narrows the length distribution mismatch by significantly increasing the candidate answer length to align with LLM outputs. Moreover, it introduces a neutral class for plausible yet under-supported candidate answers to differentiate reference incompleteness from outright incorrectness, thus enabling flexible reevaluation by consulting external knowledge for more references. Based on OBELLA, we propose a novel metric named OBELLAM (OBELLA Metric). OBELLAM integrates a cross-attention mechanism to enhance long-form candidate answer representations and employs a dynamic closed-open book evaluation strategy to tackle reference incompleteness. Our OBELLAM sets a new SOTA in aligning with human judgments across two ODQA evaluation benchmarks, marking a promising step toward more robust ODQA factuality evaluation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3e7893d5-fe9d-4bd5-9c7c-de00c37d2e1eRelated papers
- An Empirical Study of Evaluating Long-form Question AnsweringNing Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke et al.SIGIR 2025 · 2 citations
- ASQA: Factoid Questions Meet Long-Form AnswersIvan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei ChangEMNLP 2022 · 51 citations
- SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QASher Badshah, Ali Emami, Hassan SajjadACL 2026 · 1 citation
- Query Refinement Prompts for Closed-Book Long-Form QAReinald Kim Amplayo, Kellie Webster, Michael Collins, Dipanjan Das et al.ACL 2023 · 5 citations
- A Critical Evaluation of Evaluations for Long-form Question AnsweringFangyuan Xu, Yixiao Song, Mohit Iyyer, Eunsol ChoiACL 2023 · 25 citations
