QAConv: Question Answering on Informative Conversations
Chien-Sheng Wu, Andrea Madotto, Wenhao Liu, Pascale Fung, Caiming Xiong
Abstract
This paper introduces QAConv, 1 , a new question answering (QA) dataset that uses conversations as a knowledge source. We focus on informative conversations, including business emails, panel discussions, and work channels. Unlike open-domain and task-oriented dialogues, these conversations are usually long, complex, asynchronous, and involve strong domain knowledge. In total, we collect 34,608 QA pairs from 10,259 selected conversations with both human-written and machinegenerated questions. We use a question generator and a dialogue summarizer as auxiliary tools to collect and recommend questions. The dataset has two testing scenarios: chunk mode and full mode, depending on whether the grounded partial conversation is provided or retrieved. Experimental results show that stateof-the-art pretrained QA systems have limited zero-shot performance and tend to predict our questions as unanswerable. Our dataset provides a new training and evaluation testbed to facilitate QA on conversations research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b8ed5fd-70f3-4f87-93c1-98c095dc3cc4Cited by top-tier papers5
- Measuring the Effect of Transcription Noise on Downstream Language Understanding TasksOri Shapira, Shlomo E. Chazan, Amir David Nissan CohenACL 2025 · 3 citations
- Long-Tailed Question Answering in an Open WorldYi Dai, Hao Lang, Yinhe Zheng, Fei Huang et al.ACL 2023 · 3 citations
- Contrastive Learning for Inference in DialogueEtsuko Ishii, Yan Xu, Bryan Wilie, Ziwei Ji et al.EMNLP 2023 · 1 citation
- Debiased Dual-Invariant Defense for Adversarially Robust Person Re-IdentificationYuhang Zhou, Yanxiang Zhao, Zhongyun Hua, Zhipu Liu et al.AAAI 2026
- FlowRAG: Continual Learning for Dynamic Retriever in Retrieval-Augmented GenerationSenlei Zhang, Tongjun Shi, Dandan Song, Luan Zhang et al.WWW 2026
Builds on4
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- TOD-BERT: Pre-trained Natural Language Understanding for Task-Oriented DialogueChien-Sheng Wu, Steven C. H. Hoi, Richard Socher, Caiming XiongEMNLP 2020 · 210 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- DoQA - Accessing Domain-Specific FAQs via Conversational QAJon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu et al.ACL 2020 · 2 citations
Related papers
- Building and Evaluating Open-Domain Dialogue Corpora with Clarifying QuestionsMohammad Aliannejadi, Julia Kiseleva, Aleksandr Chuklin, Jeff Dalton et al.EMNLP 2021 · 61 citations
- A Dataset of Argumentative Dialogues on Scientific PapersFederico Ruggeri, Mohsen Mesgar, Iryna GurevychACL 2023
- Beyond Goldfish Memory: Long-Term Open-Domain ConversationJing Xu, Arthur Szlam, Jason WestonACL 2022 · 329 citations
- ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question AnsweringZhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma et al.EMNLP 2022 · 57 citations
- From Chat Logs to Collective Insights: Aggregative Question AnsweringWentao Zhang, Woojeong Kim, Yuntian DengEMNLP 2025
