Generating Information-Seeking Conversations from Unlabeled Documents
Gangwoo Kim, Sungdong Kim, Kang Min Yoo, Jaewoo Kang
摘要
Synthesizing datasets for conversational question answering (CQA) from unlabeled documents remains challenging due to its interactive nature.Moreover, while modeling information needs is an essential key, only few studies have discussed it.In this paper, we introduce a novel framework, SimSeek, (Simulating information-Seeking conversation from unlabeled documents), and compare its two variants.In our baseline, SimSeek-sym, a questioner generates follow-up questions upon the predetermined answer by an answerer.On the contrary, SimSeek-asym first generates the question and then finds its corresponding answer under the conversational context.Our experiments show that they can synthesize effective training resources for CQA and conversational search tasks.As a result, conversations from SimSeek-asym not only make more improvements in our experiments but also are favorably reviewed in a human evaluation.We finally release a large-scale resource of synthetic conversations, Wiki-SimSeek, containing 2 million CQA pairs built upon Wikipedia documents.With the dataset, our CQA model achieves the state-of-the-art performance on a recent CQA benchmark, QuAC.The code and dataset are available at https://github.com/naver-ai/simseek
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning to Reason and Memorize with Self-NotesJack Lanchantin, Shubham Toshniwal, Jason Weston, Arthur Szlam 等NeurIPS 2023 · 被引用 45 次
- Position: LLMs Can be Good Tutors in English EducationJingheng Ye, Shen Wang, Deqing Zou, Yibo Yan 等EMNLP 2025 · 被引用 2 次
- Reduce Human Labor On Evaluating Conversational Information Retrieval System: A Human-Machine Collaboration ApproachChen Huang, Peixin Qin, Wenqiang Lei, Jiancheng LvEMNLP 2023 · 被引用 1 次
- QUDeval: The Evaluation of Questions Under Discussion Discourse ParsingYating Wu, Ritika Mangla, Greg Durrett, Junyi Jessy LiEMNLP 2023 · 被引用 1 次
它引用的顶会 Paper9
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- Open-Retrieval Conversational Question AnsweringChen Qu, Liu Yang, Cen Chen, Minghui Qiu 等SIGIR 2020 · 被引用 84 次
- Dialog Inpainting: Turning Documents into DialogsZhuyun Dai, Arun Tejasvi Chaganty, Vincent Y. Zhao, Aida Amini 等ICML 2022 · 被引用 77 次
- FewshotQA: A simple framework for few-shot learning of question answering tasks using pre-trained text-to-text modelsRakesh Chada, Pradeep NatarajanEMNLP 2021 · 被引用 36 次
相关 Paper
- Dialogizer: Context-aware Conversational-QA Dataset Generation from Textual SourcesYerin Hwang, Yongil Kim, Hyunkyung Bae, Hwanhee Lee 等EMNLP 2023 · 被引用 3 次
- You Make me Feel like a Natural Question: Training QA Systems on Transformed Trivia QuestionsTasnim Kabir, Yoo Yeon Sung, Saptarashmi Bandyopadhyay, Hao Zou 等EMNLP 2024
- DoQA - Accessing Domain-Specific FAQs via Conversational QAJon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu 等ACL 2020 · 被引用 2 次
- LIQUID: A Framework for List Question Answering Dataset GenerationSeongyun Lee, Hyunjae Kim, Jaewoo KangAAAI 2023 · 被引用 31 次
- SocialSim: Towards Socialized Simulation of Emotional Support ConversationZhuang Chen, Yaru Cao, Guanqun Bi, Jincenzi Wu 等AAAI 2025 · 被引用 12 次
