ASQA: Factoid Questions Meet Long-Form Answers
Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei Chang
Abstract
An abundance of datasets and availability of reliable evaluation metrics have resulted in strong progress in factoid question answering (QA). This progress, however, does not easily transfer to the task of long-form QA, where the goal is to answer questions that require in-depth explanations. The hurdles include (i) a lack of high-quality data, and (ii) the absence of a well-defined notion of the answer's quality. In this work, we address these problems by (i) releasing a novel dataset and a task that we call ASQA (Answer Summaries for Questions which are Ambiguous); and (ii) proposing a reliable metric for measuring performance on ASQA. Our task focuses on factoid questions that are ambiguous, that is, have different correct answers depending on interpretation. Answers to ambiguous questions should synthesize factual information from multiple sources into a long-form summary that resolves the ambiguity. In contrast to existing long-form QA tasks (such as ELI5), ASQA admits a clear notion of correctness: a user faced with a good summary should be able to answer different interpretations of the original ambiguous question. We use this notion of correctness to define an automated metric of performance for ASQA. Our analysis demonstrates an agreement between this metric and human judgments, and reveals a considerable gap between human performance and strong baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d51008d0-9a3d-4a68-ab8a-57172d86d7b6Cited by top-tier papers85
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri et al.NeurIPS 2023 · 516 citations
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun et al.EMNLP 2023 · 315 citations
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill SetsSeonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang et al.ICLR 2024 · 176 citations
- Enabling Large Language Models to Generate Text with CitationsTianyu Gao, Howard Yen, Jiatong Yu, Danqi ChenEMNLP 2023 · 152 citations
Builds on9
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 317 citations
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 162 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler et al.NeurIPS 2020 · 124 citations
Related papers
- An Empirical Study of Evaluating Long-form Question AnsweringNing Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke et al.SIGIR 2025 · 2 citations
- Concise Answers to Complex Questions: Summarization of Long-form AnswersAbhilash Potluri, Fangyuan Xu, Eunsol ChoiACL 2023 · 4 citations
- Query Refinement Prompts for Closed-Book Long-Form QAReinald Kim Amplayo, Kellie Webster, Michael Collins, Dipanjan Das et al.ACL 2023 · 5 citations
- OBELLA: Open the Book for Evaluating Long-Form Large Language Model Answers in Open-Domain Question AnsweringTianyu Ren, Zhaoyu Zhang, Hui Wang, Karen RaffertySIGIR 2025
- LFQA-E: Carefully Benchmarking Long-form QA EvaluationYuchen Fan, Chen Ling, Xin Zhong, Shuo Zhang et al.ICLR 2026 · 2 citations
