DialoGPS: Dialogue Path Sampling in Continuous Semantic Space for Data Augmentation in Multi-Turn Conversations
Ang Lv, Jinpeng Li, Yuhan Chen, Gao Xing, Ji Zhang, Rui Yan
Abstract
In open-domain dialogue generation tasks, contexts and responses in most datasets are one-to-one mapped, violating an important many-to-many characteristic: a context leads to various responses, and a response answers multiple contexts. Without such patterns, models poorly generalize and prefer responding safely. Many attempts have been made in either multi-turn settings from a one-to-many perspective or in a many-to-many perspective but limited to single-turn settings. The major challenge to many-to-many augment multi-turn dialogues is that discretely replacing each turn with semantic similarity breaks fragile context coherence. In this paper, we propose DialoGue Path Sampling (DialoGPS) method in continuous semantic space, the first many-to-many augmentation method for multi-turn dialogues. Specifically, we map a dialogue to our extended Brownian Bridge, a special Gaussian process. We sample latent variables to form coherent dialogue paths in the continuous space. A dialogue path corresponds to a new multi-turn dialogue and is used as augmented training data. We show the effect of DialoGPS with both automatic and human evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f29076d0-a008-4a31-9d8d-0c1260db03deCited by top-tier papers4
- StreamingDialogue: Prolonged Dialogue Learning via Long Context Compression with Minimal LossesJianan Li, Quan Tu, Cunli Mao, Zhengtao Yu et al.NeurIPS 2024 · 13 citations
- "In-Dialogues We Learn": Towards Personalized Dialogue Without Pre-defined Profiles through In-Dialogue LearningChuanqi Cheng, Quan Tu, Wei Wu, Shuo Shang et al.EMNLP 2024 · 3 citations
- Explicit and Implicit Data Augmentation for Social Event DetectionCongbo Ma, Yuxia Wang, Jia Wu, Jian Yang et al.ACL 2025 · 2 citations
- EventWeave: A Dynamic Framework for Capturing Core and Supporting Events in Dialogue SystemsZhengyi Zhao, Shubo Zhang, Yiming Du, Bin Liang et al.ACL 2026 · 2 citations
Builds on10
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Knowledge-Grounded Dialogue Generation with Pre-trained Language ModelsXueliang Zhao, Wei Wu, Can Xu, Chongyang Tao et al.EMNLP 2020 · 153 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- Language modeling via stochastic processesRose E. Wang, Esin Durmus, Noah D. Goodman, Tatsunori HashimotoICLR 2022 · 28 citations
Related papers
- Generating Dialogue Responses from a Semantic Latent SpaceWei-Jen Ko, Avik Ray, Yilin Shen, Hongxia JinEMNLP 2020 · 4 citations
- Counterfactual Data Augmentation via Perspective Transition for Open-Domain DialoguesJiao Ou, Jinchao Zhang, Yang Feng, Jie ZhouEMNLP 2022 · 9 citations
- DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response GenerationWei Chen, Yeyun Gong, Song Wang, Bolun Yao et al.ACL 2022
- Re³Dial: Retrieve, Reorganize and Rescale Conversations for Long-Turn Open-Domain Dialogue Pre-trainingJiaxin Wen, Hao Zhou, Jian Guan, Jie Zhou et al.EMNLP 2023 · 2 citations
- Task-Oriented Dialog Systems That Consider Multiple Appropriate Responses under the Same ContextYichi Zhang, Zhijian Ou, Zhou YuAAAI 2020 · 198 citations
