Ditch the Gold Standard: Re-evaluating Conversational Question Answering
Huihan Li, Tianyu Gao, Manan Goenka, Danqi Chen
Abstract
Conversational question answering aims to provide natural-language answers to users in information-seeking conversations. Existing conversational QA benchmarks compare models with pre-collected human-human conversations, using ground-truth answers provided in conversational history. It remains unclear whether we can rely on this static evaluation for model development and whether current systems can well generalize to real-world human-machine conversations. In this work, we conduct the first large-scale human evaluation of state-of-the-art conversational QA systems, where human evaluators converse with models and judge the correctness of their answers. We find that the distribution of human machine conversations differs drastically from that of human-human conversations, and there is a disagreement between human and gold-history evaluation in terms of model ranking. We further investigate how to improve automatic evaluations, and propose a question rewriting mechanism based on predicted history, which better correlates with human judgments. Finally, we analyze the impact of various modeling strategies and discuss future directions towards building better conversational question answering systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64d0b319-9b56-47bd-ba2b-70304e044c0dCited by top-tier papers5
- Dialog Inpainting: Turning Documents into DialogsZhuyun Dai, Arun Tejasvi Chaganty, Vincent Y. Zhao, Aida Amini et al.ICML 2022 · 77 citations
- Learning to Relate to Previous Turns in Conversational SearchFengran Mo, Jian-Yun Nie, Kaiyu Huang, Kelong Mao et al.KDD 2023 · 16 citations
- ChatRetriever: Adapting Large Language Models for Generalized and Robust Conversational Dense RetrievalKelong Mao, Chenlong Deng, Haonan Chen, Fengran Mo et al.EMNLP 2024 · 7 citations
- Generating Information-Seeking Conversations from Unlabeled DocumentsGangwoo Kim, Sungdong Kim, Kang Min Yoo, Jaewoo KangEMNLP 2022 · 4 citations
- Reduce Human Labor On Evaluating Conversational Information Retrieval System: A Human-Machine Collaboration ApproachChen Huang, Peixin Qin, Wenqiang Lei, Jiancheng LvEMNLP 2023 · 1 citation
Builds on6
- Open-Retrieval Conversational Question AnsweringChen Qu, Liu Yang, Cen Chen, Minghui Qiu et al.SIGIR 2020 · 84 citations
- GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue SystemsLishan Huang, Zheng Ye, Jinghui Qin, Liang Lin et al.EMNLP 2020 · 73 citations
- USR: An Unsupervised and Reference Free Evaluation Metric for Dialog GenerationShikib Mehri, Maxine EskénaziACL 2020 · 10 citations
- DoQA - Accessing Domain-Specific FAQs via Conversational QAJon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu et al.ACL 2020 · 2 citations
- Towards Quantifiable Dialogue Coherence EvaluationZheng Ye, Liucun Lu, Lishan Huang, Liang Lin et al.ACL 2021
Related papers
- Interview Evaluation: A Novel Approach for Automatic Evaluation of Conversational Question Answering ModelsXibo Li, Bowei Zou, Yifan Fan, Yanling Li et al.EMNLP 2023
- CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement LearningZeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter et al.EMNLP 2022 · 37 citations
- An Empirical Study of Evaluating Long-form Question AnsweringNing Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke et al.SIGIR 2025 · 2 citations
- CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational SearchFengran Mo, Abbas Ghaddar, Kelong Mao, Mehdi Rezagholizadeh et al.EMNLP 2024 · 9 citations
- ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question AnsweringSimon Lupart, Mohammad Aliannejadi, Evangelos KanoulasACL 2026 · 5 citations
