A Synthetic Data Generation Framework for Grounded Dialogues
Jianzhu Bao, Rui Wang, Yasheng Wang, Aixin Sun, Yitong Li, Fei Mi, Ruifeng Xu
摘要
Training grounded response generation models often requires a large collection of grounded dialogues. However, it is costly to build such dialogues. In this paper, we present a synthetic data generation framework (SynDG) for grounded dialogues. The generation process utilizes large pre-trained language models and freely available knowledge data (e.g., Wikipedia pages, persona profiles, etc.). The key idea of designing SynDG is to consider dialogue flow and coherence in the generation process. Specifically, given knowledge data, we first heuristically determine a dialogue flow, which is a series of knowledge pieces. Then, we employ T5 to incrementally turn the dialogue flow into a dialogue. To ensure coherence of both the dialogue flow and the synthetic dialogue, we design a two-level filtering strategy, at the flow-level and the utterance-level respectively. Experiments on two public benchmarks show that the synthetic grounded dialogue data produced by our framework is able to significantly boost model performance in both full training data and low-resource scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Simulacrum of Stories: Examining Large Language Models as Qualitative Research ParticipantsShivani Kapania, William Agnew, Motahhare Eslami, Hoda Heidari 等CHI 2025 · 被引用 59 次
- Words Like Knives: Backstory-Personalized Modeling and Detection of Violent CommunicationJocelyn J. Shen, Akhila Yerukola, Xuhui Zhou, Cynthia Breazeal 等EMNLP 2025 · 被引用 4 次
- Inductive-Deductive Strategy Reuse for Multi-Turn Instructional DialoguesJiao Ou, Jiayu Wu, Che Liu, Fuzheng Zhang 等EMNLP 2024 · 被引用 2 次
- Deep-Reporter: Deep Research for Grounded Multimodal Long-Form GenerationFangda Ye, Kuicai Dong, Zhifei Xie, Yuxin Hu 等ACL 2026 · 被引用 2 次
- ArgGenBench: Benchmarking the Complex Controlled Argument Generation Capability of Large Language ModelsBojun Jin, Jianzhu Bao, Yang Sun, Yice Zhang 等ACL 2026
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Sequential Latent Knowledge Selection for Knowledge-Grounded DialogueByeongchang Kim, Jaewoo Ahn, Gunhee KimICLR 2020 · 被引用 179 次
- Bridging the Gap between Prior and Posterior Knowledge Selection for Knowledge-Grounded Dialogue GenerationXiuyi Chen, Fandong Meng, Peng Li, Feilong Chen 等EMNLP 2020 · 被引用 78 次
- Dialog Inpainting: Turning Documents into DialogsZhuyun Dai, Arun Tejasvi Chaganty, Vincent Y. Zhao, Aida Amini 等ICML 2022 · 被引用 77 次
- Zero-Resource Knowledge-Grounded Dialogue GenerationLinxiao Li, Can Xu, Wei Wu, Yufan Zhao 等NeurIPS 2020 · 被引用 75 次
相关 Paper
- KPT: Keyword-Guided Pre-training for Grounded Dialog GenerationQi Zhu, Fei Mi, Zheng Zhang, Yasheng Wang 等AAAI 2023 · 被引用 5 次
- Knowledge-Grounded Dialogue Generation with Pre-trained Language ModelsXueliang Zhao, Wei Wu, Can Xu, Chongyang Tao 等EMNLP 2020 · 被引用 153 次
- Low-Resource Knowledge-Grounded Dialogue GenerationXueliang Zhao, Wei Wu, Chongyang Tao, Can Xu 等ICLR 2020 · 被引用 115 次
- RetGen: A Joint Framework for Retrieval and Grounded Text Generation ModelingYizhe Zhang, Siqi Sun, Xiang Gao, Yuwei Fang 等AAAI 2022 · 被引用 45 次
- Call for Customized Conversation: Customized Conversation Grounding Persona and KnowledgeYoonna Jang, Jungwoo Lim, Yuna Hur, Dongsuk Oh 等AAAI 2022 · 被引用 47 次
