Counterfactual Data Augmentation via Perspective Transition for Open-Domain Dialogues
Jiao Ou, Jinchao Zhang, Yang Feng, Jie Zhou
摘要
The construction of open-domain dialogue systems requires high-quality dialogue datasets. The dialogue data admits a wide variety of responses for a given dialogue history, especially responses with different semantics. However, collecting high-quality such a dataset in most scenarios is labor-intensive and time-consuming. In this paper, we propose a data augmentation method to automatically augment high-quality responses with different semantics by counterfactual inference. Specifically, given an observed dialogue, our counterfactual generation model first infers semantically different responses by replacing the observed reply perspective with substituted ones. Furthermore, our data selection method filters out detrimental augmented responses. Experimental results show that our data augmentation method can augment high-quality responses with different semantics for a given dialogue history, and can outperform competitive baselines on multiple downstream tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SODA: Million-scale Dialogue Distillation with Social Commonsense ContextualizationHyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West 等EMNLP 2023 · 被引用 60 次
- Improving Conversational Recommendation Systems via Counterfactual Data SimulationXiaolei Wang, Kun Zhou, Xinyu Tang, Wayne Xin Zhao 等KDD 2023 · 被引用 12 次
- Inductive-Deductive Strategy Reuse for Multi-Turn Instructional DialoguesJiao Ou, Jiayu Wu, Che Liu, Fuzheng Zhang 等EMNLP 2024 · 被引用 2 次
- A Causal Approach for Counterfactual Reasoning in NarrativesFeiteng Mu, Wenjie LiACL 2024 · 被引用 1 次
- CHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme IngredientsJaehyung Seo, Hyeonseok Moon, Jaewook Lee, Sugyeong Eo 等EMNLP 2023 · 被引用 1 次
它引用的顶会 Paper14
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers 等ICML 2020 · 被引用 242 次
- Data Manipulation: Towards Effective Instance Learning for Neural Dialogue Generation via Learning to Augment and ReweightHengyi Cai, Hongshen Chen, Yonghao Song, Cheng Zhang 等ACL 2020 · 被引用 57 次
相关 Paper
- Counterfactual Off-Policy Training for Neural Dialogue GenerationQingfu Zhu, Wei-Nan Zhang, Ting Liu, William Yang WangEMNLP 2020 · 被引用 18 次
- Dialogue Distillation: Open-Domain Dialogue Augmentation Using Unpaired DataRongsheng Zhang, Yinhe Zheng, Jianzhi Shao, Xiaoxi Mao 等EMNLP 2020 · 被引用 25 次
- Learning towards Selective Data Augmentation for Dialogue GenerationXiuying Chen, Mingzhe Li, Jiayi Zhang, Xiaoqiang Xia 等AAAI 2023 · 被引用 7 次
- Retrieval-guided Counterfactual Generation for QABhargavi Paranjape, Matthew Lamm, Ian TenneyACL 2022 · 被引用 39 次
- Exploring the Efficacy of Automatically Generated Counterfactuals for Sentiment AnalysisLinyi Yang, Jiazheng Li, Padraig Cunningham, Yue Zhang 等ACL 2021
