Dialog Inpainting: Turning Documents into Dialogs
Zhuyun Dai, Arun Tejasvi Chaganty, Vincent Y. Zhao, Aida Amini, Qazi Mamunur Rashid, Mike Green, Kelvin Guu
摘要
Many important questions (e.g."How to eat healthier?") require conversation to establish context and explore in depth. However, conversational question answering (ConvQA) systems have long been stymied by scarce training data that is expensive to collect. To address this problem, we propose a new technique for synthetically generating diverse and high-quality dialog data: dialog inpainting. Our approach takes the text of any document and transforms it into a two-person dialog between the writer and an imagined reader: we treat sentences from the article as utterances spoken by the writer, and then use a dialog inpainter to predict what the imagined reader asked or said in between each of the writer's utterances. By applying this approach to passages from Wikipedia and the web, we produce WikiDialog and WebDialog, two datasets totalling 19 million diverse information-seeking dialogs -- 1,000x larger than the largest existing ConvQA dataset. Furthermore, human raters judge the answer adequacy and conversationality of WikiDialog to be as good or better than existing manually-collected datasets. Using our inpainted data to pre-train ConvQA retrieval systems, we significantly advance state-of-the-art across three benchmarks (QReCC, OR-QuAC, TREC CAsT) yielding up to 40% relative gains on standard evaluation metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- PEER: A Collaborative Language ModelTimo Schick, Jane A. Yu, Zhengbao Jiang, Fabio Petroni 等ICLR 2023 · 被引用 44 次
- ConvTrans: Transforming Web Search Sessions for Conversational Dense RetrievalKelong Mao, Zhicheng Dou, Hongjin Qian, Fengran Mo 等EMNLP 2022 · 被引用 12 次
- A Synthetic Data Generation Framework for Grounded DialoguesJianzhu Bao, Rui Wang, Yasheng Wang, Aixin Sun 等ACL 2023 · 被引用 11 次
- CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational SearchFengran Mo, Abbas Ghaddar, Kelong Mao, Mehdi Rezagholizadeh 等EMNLP 2024 · 被引用 9 次
- Achieving Human Parity in Content-Grounded Datasets GenerationAsaf Yehudai, Boaz Carmeli, Yosi Mass, Ofir Arviv 等ICLR 2024 · 被引用 9 次
它引用的顶会 Paper9
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- Pre-training Tasks for Embedding-based Large-scale RetrievalWei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang 等ICLR 2020 · 被引用 325 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- doc2dial: A Goal-Oriented Document-Grounded Dialogue DatasetSong Feng, Hui Wan, R. Chulaka Gunasekara, Siva Sankalp Patel 等EMNLP 2020 · 被引用 87 次
- Open-Retrieval Conversational Question AnsweringChen Qu, Liu Yang, Cen Chen, Minghui Qiu 等SIGIR 2020 · 被引用 84 次
相关 Paper
- Dialogizer: Context-aware Conversational-QA Dataset Generation from Textual SourcesYerin Hwang, Yongil Kim, Hyunkyung Bae, Hwanhee Lee 等EMNLP 2023 · 被引用 3 次
- Generating Information-Seeking Conversations from Unlabeled DocumentsGangwoo Kim, Sungdong Kim, Kang Min Yoo, Jaewoo KangEMNLP 2022 · 被引用 4 次
- q2d: Turning Questions into Dialogs to Teach Models How to SearchYonatan Bitton, Shlomi Cohen-Ganor, Ido Hakimi, Yoad Lewenberg 等EMNLP 2023 · 被引用 3 次
- DoQA - Accessing Domain-Specific FAQs via Conversational QAJon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu 等ACL 2020 · 被引用 2 次
- CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement LearningZeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter 等EMNLP 2022 · 被引用 37 次
