A Critical Evaluation of Evaluations for Long-form Question Answering
Fangyuan Xu, Yixiao Song, Mohit Iyyer, Eunsol Choi
摘要
Long-form question answering (LFQA) enables answering a wide range of questions, but its flexibility poses enormous challenges for evaluation. We perform the first targeted study of the evaluation of long-form answers, covering both human and automatic evaluation practices. We hire domain experts in seven areas to provide preference judgments over pairs of answers, along with free-form justifications for their choices. We present a careful analysis of experts' evaluation, which focuses on new aspects such as the comprehensiveness of the answer. Next, we examine automatic text generation metrics, finding that no existing metrics are predictive of human preference judgments. However, some metrics correlate with fine-grained aspects of answers (e.g., coherence). We encourage future work to move away from a single "overall score" of the answer and adopt a multi-faceted evaluation, targeting aspects such as factuality and completeness. We publicly release all of our annotations and code to spur future work into LFQA evaluation. 1 * * Equal contribution. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis 等EMNLP 2023 · 被引用 225 次
- Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and InconsistenciesSunnie S. Y. Kim, Jennifer Wortman Vaughan, Q. Vera Liao, Tania Lombrozo 等CHI 2025 · 被引用 118 次
- Human Feedback is not Gold StandardTom Hosking, Phil Blunsom, Max BartoloICLR 2024 · 被引用 96 次
- Reinforcement Learning with Evolving Rubrics for Deep ResearchRulin Shao, Akari Asai, Shannon Shen, Hamish Ivison 等ICML 2026 · 被引用 78 次
- QASA: Advanced Question Answering on Scientific ArticlesYoonjoo Lee, Kyungjae Lee, Sunghyun Park, Dasol Hwang 等ICML 2023 · 被引用 76 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- An Empirical Study of Evaluating Long-form Question AnsweringNing Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke 等SIGIR 2025 · 被引用 2 次
- LFQA-E: Carefully Benchmarking Long-form QA EvaluationYuchen Fan, Chen Ling, Xin Zhong, Shuo Zhang 等ICLR 2026 · 被引用 2 次
- Query Refinement Prompts for Closed-Book Long-Form QAReinald Kim Amplayo, Kellie Webster, Michael Collins, Dipanjan Das 等ACL 2023 · 被引用 5 次
- OBELLA: Open the Book for Evaluating Long-Form Large Language Model Answers in Open-Domain Question AnsweringTianyu Ren, Zhaoyu Zhang, Hui Wang, Karen RaffertySIGIR 2025
- ASQA: Factoid Questions Meet Long-Form AnswersIvan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei ChangEMNLP 2022 · 被引用 51 次
