Generating Information-Seeking Conversations from Unlabeled Documents
Gangwoo Kim, Sungdong Kim, Kang Min Yoo, Jaewoo Kang
Abstract
Synthesizing datasets for conversational question answering (CQA) from unlabeled documents remains challenging due to its interactive nature.Moreover, while modeling information needs is an essential key, only few studies have discussed it.In this paper, we introduce a novel framework, SimSeek, (Simulating information-Seeking conversation from unlabeled documents), and compare its two variants.In our baseline, SimSeek-sym, a questioner generates follow-up questions upon the predetermined answer by an answerer.On the contrary, SimSeek-asym first generates the question and then finds its corresponding answer under the conversational context.Our experiments show that they can synthesize effective training resources for CQA and conversational search tasks.As a result, conversations from SimSeek-asym not only make more improvements in our experiments but also are favorably reviewed in a human evaluation.We finally release a large-scale resource of synthetic conversations, Wiki-SimSeek, containing 2 million CQA pairs built upon Wikipedia documents.With the dataset, our CQA model achieves the state-of-the-art performance on a recent CQA benchmark, QuAC.The code and dataset are available at https://github.com/naver-ai/simseek
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2682474b-5ed7-4d2a-8024-16216510694cCited by top-tier papers4
- Learning to Reason and Memorize with Self-NotesJack Lanchantin, Shubham Toshniwal, Jason Weston, Arthur Szlam et al.NeurIPS 2023 · 45 citations
- Position: LLMs Can be Good Tutors in English EducationJingheng Ye, Shen Wang, Deqing Zou, Yibo Yan et al.EMNLP 2025 · 2 citations
- Reduce Human Labor On Evaluating Conversational Information Retrieval System: A Human-Machine Collaboration ApproachChen Huang, Peixin Qin, Wenqiang Lei, Jiancheng LvEMNLP 2023 · 1 citation
- QUDeval: The Evaluation of Questions Under Discussion Discourse ParsingYating Wu, Ritika Mangla, Greg Durrett, Junyi Jessy LiEMNLP 2023 · 1 citation
Builds on9
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Open-Retrieval Conversational Question AnsweringChen Qu, Liu Yang, Cen Chen, Minghui Qiu et al.SIGIR 2020 · 84 citations
- Dialog Inpainting: Turning Documents into DialogsZhuyun Dai, Arun Tejasvi Chaganty, Vincent Y. Zhao, Aida Amini et al.ICML 2022 · 77 citations
- FewshotQA: A simple framework for few-shot learning of question answering tasks using pre-trained text-to-text modelsRakesh Chada, Pradeep NatarajanEMNLP 2021 · 36 citations
Related papers
- Dialogizer: Context-aware Conversational-QA Dataset Generation from Textual SourcesYerin Hwang, Yongil Kim, Hyunkyung Bae, Hwanhee Lee et al.EMNLP 2023 · 3 citations
- You Make me Feel like a Natural Question: Training QA Systems on Transformed Trivia QuestionsTasnim Kabir, Yoo Yeon Sung, Saptarashmi Bandyopadhyay, Hao Zou et al.EMNLP 2024
- DoQA - Accessing Domain-Specific FAQs via Conversational QAJon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu et al.ACL 2020 · 2 citations
- LIQUID: A Framework for List Question Answering Dataset GenerationSeongyun Lee, Hyunjae Kim, Jaewoo KangAAAI 2023 · 31 citations
- SocialSim: Towards Socialized Simulation of Emotional Support ConversationZhuang Chen, Yaru Cao, Guanqun Bi, Jincenzi Wu et al.AAAI 2025 · 12 citations
