Dialogizer: Context-aware Conversational-QA Dataset Generation from Textual Sources
Yerin Hwang, Yongil Kim, Hyunkyung Bae, Hwanhee Lee, Jeesoo Bang, Kyomin Jung
摘要
To address the data scarcity issue in Conversational question answering (ConvQA), a dialog inpainting method, which utilizes documents to generate ConvQA datasets, has been proposed. However, the original dialog inpainting model is trained solely on the dialog reconstruction task, resulting in the generation of questions with low contextual relevance due to insufficient learning of question-answer alignment. To overcome this limitation, we propose a novel framework called Dialogizer, which has the capability to automatically generate ConvQA datasets with high contextual relevance from textual sources. The framework incorporates two training tasks: question-answer matching (QAM) and topic-aware dialog generation (TDG). Moreover, re-ranking is conducted during the inference phase based on the contextual relevance of the generated questions. Using our framework, we produce four Con-vQA datasets by utilizing documents from multiple domains as the primary source. Through automatic evaluation using diverse metrics, as well as human evaluation, we validate that our proposed framework exhibits the ability to generate datasets of higher quality compared to the baseline dialog inpainting model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Open-Retrieval Conversational Question AnsweringChen Qu, Liu Yang, Cen Chen, Minghui Qiu 等SIGIR 2020 · 被引用 84 次
- Dialog Inpainting: Turning Documents into DialogsZhuyun Dai, Arun Tejasvi Chaganty, Vincent Y. Zhao, Aida Amini 等ICML 2022 · 被引用 77 次
- GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue SystemsLishan Huang, Zheng Ye, Jinghui Qin, Liang Lin 等EMNLP 2020 · 被引用 73 次
- Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationBo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou 等ACL 2020 · 被引用 69 次
- Dialogue Response Ranking Training with Large-Scale Human Feedback DataXiang Gao, Yizhe Zhang, Michel Galley, Chris Brockett 等EMNLP 2020 · 被引用 67 次
相关 Paper
- Generating Information-Seeking Conversations from Unlabeled DocumentsGangwoo Kim, Sungdong Kim, Kang Min Yoo, Jaewoo KangEMNLP 2022 · 被引用 4 次
- CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement LearningZeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter 等EMNLP 2022 · 被引用 37 次
- Contrastive Domain Adaptation for Question Answering using Limited Text CorporaZhenrui Yue, Bernhard Kratzwald, Stefan FeuerriegelEMNLP 2021 · 被引用 22 次
- Unsupervised and Pseudo-Supervised Vision-Language Alignment in Visual DialogFeilong Chen, Duzhen Zhang, Xiuyi Chen, Jing Shi 等ACM MM 2022 · 被引用 10 次
- TARGA: Targeted Synthetic Data Generation for Practical Reasoning over Structured DataXiang Huang, Jiayu Shen, Shanshan Huang, Sitao Cheng 等ACL 2025
