How Do We Answer Complex Questions: Discourse Structure of Long-form Answers
Fangyuan Xu, Junyi Jessy Li, Eunsol Choi
摘要
Long-form answers, consisting of multiple sentences, can provide nuanced and comprehensive answers to a broader set of questions. To better understand this complex and understudied task, we study the functional structure of long-form answers collected from three datasets, ELI5 (Fan et al., 2019) , We-bGPT (Nakano et al., 2021) and Natural Questions (Kwiatkowski et al., 2019) . Our main goal is to understand how humans organize information to craft complex answers. We develop an ontology of six sentence-level functional roles for long-form answers, and annotate 3.9k sentences in 640 answer paragraphs. Different answer collection methods manifest in different discourse structures. We further analyze model-generated answers -finding that annotators agree less with each other when annotating model-generated answers compared to annotating human-written answers. Our annotated data enables training a strong classifier that can be used for automatic analysis. We hope our work can inspire future research on discourse-level modeling and evaluation of long-form QA systems. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill SetsSeonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang 等ICLR 2024 · 被引用 176 次
- QASA: Advanced Question Answering on Scientific ArticlesYoonjoo Lee, Kyungjae Lee, Sunghyun Park, Dasol Hwang 等ICML 2023 · 被引用 76 次
- WebCPM: Interactive Web Search for Chinese Long-form Question AnsweringYujia Qin, Zihan Cai, Dian Jin, Lan Yan 等ACL 2023 · 被引用 25 次
- Foundational Autoraters: Taming Large Language Models for Better Automatic EvaluationTu Vu, Kalpesh Krishna, Salaheddin Alzubi, Chris Tar 等EMNLP 2024 · 被引用 14 次
- CREPE: Open-Domain Question Answering with False PresuppositionsXinyan Yu, Sewon Min, Luke Zettlemoyer, Hannaneh HajishirziACL 2023 · 被引用 13 次
它引用的顶会 Paper4
- Discourse as a Function of Event: Profiling Discourse Structure in News Articles around the Main EventPrafulla Kumar Choubey, Aaron Lee, Ruihong Huang, Lu WangACL 2020 · 被引用 54 次
- The Perils of Using Mechanical Turk to Evaluate Open-Ended Text GenerationMarzena Karpinska, Nader Akoury, Mohit IyyerEMNLP 2021 · 被引用 3 次
- Which Linguist Invented the Lightbulb? Presupposition Verification for Question-AnsweringNajoung Kim, Ellie Pavlick, Burcu Karagol Ayan, Deepak RamachandranACL 2021
- Controllable Open-ended Question Generation with A New Question Type OntologyShuyang Cao, Lu WangACL 2021
相关 Paper
- Concise Answers to Complex Questions: Summarization of Long-form AnswersAbhilash Potluri, Fangyuan Xu, Eunsol ChoiACL 2023 · 被引用 4 次
- ASQA: Factoid Questions Meet Long-Form AnswersIvan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei ChangEMNLP 2022 · 被引用 51 次
- LFQA-E: Carefully Benchmarking Long-form QA EvaluationYuchen Fan, Chen Ling, Xin Zhong, Shuo Zhang 等ICLR 2026 · 被引用 2 次
- An Empirical Study of Evaluating Long-form Question AnsweringNing Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke 等SIGIR 2025 · 被引用 2 次
- A Critical Evaluation of Evaluations for Long-form Question AnsweringFangyuan Xu, Yixiao Song, Mohit Iyyer, Eunsol ChoiACL 2023 · 被引用 25 次
