Diversifying Dialog Generation via Adaptive Label Smoothing
Yida Wang, Yinhe Zheng, Yong Jiang, Minlie Huang
Abstract
Neural dialogue generation models trained with the one-hot target distribution suffer from the over-confidence issue, which leads to poor generation diversity as widely reported in the literature. Although existing approaches such as label smoothing can alleviate this issue, they fail to adapt to diverse dialog contexts. In this paper, we propose an Adaptive Label Smoothing (AdaLabel) approach that can adaptively estimate a target label distribution at each time step for different contexts. The maximum probability in the predicted distribution is used to modify the soft target distribution produced by a novel light-weight bi-directional decoder module. The resulting target distribution is aware of both previous and future contexts and is adjusted to avoid over-training the dialogue model. Our model can be trained in an endto-end manner. Extensive experiments on two benchmark datasets show that our approach outperforms various competitive baselines in producing diverse responses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1d0154f-08e8-4f6a-9513-e5b28b96cf83Cited by top-tier papers7
- Teacher Forcing Recovers Reward Functions for Text GenerationYongchang Hao, Yuxin Liu, Lili MouNeurIPS 2022 · 24 citations
- Estimating Soft Labels for Out-of-Domain Intent DetectionHao Lang, Yinhe Zheng, Jian Sun, Fei Huang et al.EMNLP 2022 · 12 citations
- Prompt Conditioned VAE: Enhancing Generative Replay for Lifelong Learning in Task-Oriented DialogueYingxiu Zhao, Yinhe Zheng, Zhiliang Tian, Chang Gao et al.EMNLP 2022 · 7 citations
- SimOAP: Improve Coherence and Consistency in Persona-based Dialogue Generation via Over-sampling and Post-evaluationJunkai Zhou, Liang Pang, Huawei Shen, Xueqi ChengACL 2023 · 6 citations
- An Equal-Size Hard EM Algorithm for Diverse Dialogue GenerationYuqiao Wen, Yongchang Hao, Yanshuai Cao, Lili MouICLR 2023 · 2 citations
Builds on11
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- A Pre-Training Based Personalized Dialogue Generation Model with Persona-Sparse DataYinhe Zheng, Rongsheng Zhang, Minlie Huang, Xiaoxi MaoAAAI 2020 · 173 citations
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross et al.ICLR 2021 · 154 citations
Related papers
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent VariableSiqi Bao, Huang He, Fan Wang, Hua Wu et al.ACL 2020 · 229 citations
- AvgOut: A Simple Output-Probability Measure to Eliminate Dull ResponsesTong Niu, Mohit BansalAAAI 2020 · 3 citations
- Adaptive Bridge between Training and Inference for Dialogue GenerationHaoran Xu, Hainan Zhang, Yanyan Zou, Hongshen Chen et al.EMNLP 2021 · 5 citations
- Learning from Perturbations: Diverse and Informative Dialogue Generation with Inverse Adversarial TrainingWangchunshu Zhou, Qifei Li, Chenle LiACL 2021
- Diversifying Dialogue Generation with Non-Conversational TextHui Su, Xiaoyu Shen, Sanqiang Zhao, Xiao Zhou et al.ACL 2020 · 39 citations
