Ditch the Gold Standard: Re-evaluating Conversational Question Answering
Huihan Li, Tianyu Gao, Manan Goenka, Danqi Chen
摘要
Conversational question answering aims to provide natural-language answers to users in information-seeking conversations. Existing conversational QA benchmarks compare models with pre-collected human-human conversations, using ground-truth answers provided in conversational history. It remains unclear whether we can rely on this static evaluation for model development and whether current systems can well generalize to real-world human-machine conversations. In this work, we conduct the first large-scale human evaluation of state-of-the-art conversational QA systems, where human evaluators converse with models and judge the correctness of their answers. We find that the distribution of human machine conversations differs drastically from that of human-human conversations, and there is a disagreement between human and gold-history evaluation in terms of model ranking. We further investigate how to improve automatic evaluations, and propose a question rewriting mechanism based on predicted history, which better correlates with human judgments. Finally, we analyze the impact of various modeling strategies and discuss future directions towards building better conversational question answering systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Dialog Inpainting: Turning Documents into DialogsZhuyun Dai, Arun Tejasvi Chaganty, Vincent Y. Zhao, Aida Amini 等ICML 2022 · 被引用 77 次
- Learning to Relate to Previous Turns in Conversational SearchFengran Mo, Jian-Yun Nie, Kaiyu Huang, Kelong Mao 等KDD 2023 · 被引用 16 次
- ChatRetriever: Adapting Large Language Models for Generalized and Robust Conversational Dense RetrievalKelong Mao, Chenlong Deng, Haonan Chen, Fengran Mo 等EMNLP 2024 · 被引用 7 次
- Generating Information-Seeking Conversations from Unlabeled DocumentsGangwoo Kim, Sungdong Kim, Kang Min Yoo, Jaewoo KangEMNLP 2022 · 被引用 4 次
- Reduce Human Labor On Evaluating Conversational Information Retrieval System: A Human-Machine Collaboration ApproachChen Huang, Peixin Qin, Wenqiang Lei, Jiancheng LvEMNLP 2023 · 被引用 1 次
它引用的顶会 Paper6
- Open-Retrieval Conversational Question AnsweringChen Qu, Liu Yang, Cen Chen, Minghui Qiu 等SIGIR 2020 · 被引用 84 次
- GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue SystemsLishan Huang, Zheng Ye, Jinghui Qin, Liang Lin 等EMNLP 2020 · 被引用 73 次
- USR: An Unsupervised and Reference Free Evaluation Metric for Dialog GenerationShikib Mehri, Maxine EskénaziACL 2020 · 被引用 10 次
- DoQA - Accessing Domain-Specific FAQs via Conversational QAJon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu 等ACL 2020 · 被引用 2 次
- Towards Quantifiable Dialogue Coherence EvaluationZheng Ye, Liucun Lu, Lishan Huang, Liang Lin 等ACL 2021
相关 Paper
- Interview Evaluation: A Novel Approach for Automatic Evaluation of Conversational Question Answering ModelsXibo Li, Bowei Zou, Yifan Fan, Yanling Li 等EMNLP 2023
- CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement LearningZeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter 等EMNLP 2022 · 被引用 37 次
- An Empirical Study of Evaluating Long-form Question AnsweringNing Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke 等SIGIR 2025 · 被引用 2 次
- CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational SearchFengran Mo, Abbas Ghaddar, Kelong Mao, Mehdi Rezagholizadeh 等EMNLP 2024 · 被引用 9 次
- ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question AnsweringSimon Lupart, Mohammad Aliannejadi, Evangelos KanoulasACL 2026 · 被引用 5 次
