Heterogeneous-Branch Collaborative Learning for Dialogue Generation
Yiwei Li, Shaoxiong Feng, Bin Sun, Kan Li
Abstract
With the development of deep learning, advanced dialogue generation methods usually require a greater amount of computational resources. One promising approach to obtaining a high-performance and lightweight model is knowledge distillation, which relies heavily on the pre-trained powerful teacher. Collaborative learning, also known as online knowledge distillation, is an effective way to conduct one-stage group distillation in the absence of a well-trained large teacher model. However, previous work has a severe branch homogeneity problem due to the same training objective and the independent identical training sets. To alleviate this problem, we consider the dialogue attributes in the training of network branches. Each branch learns the attribute-related features based on the selected subset. Furthermore, we propose a dual group-based knowledge distillation method, consisting of positive distillation and negative distillation, to further diversify the features of different branches in a steadily and interpretable way. The proposed approach significantly improves branch heterogeneity and outperforms state-of-the-art collaborative learning methods on two widely used open-domain dialogue datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c29a1676-68f4-4c46-817e-7a05f4f59068Cited by top-tier papers3
- Turning Dust into Gold: Distilling Complex Reasoning Capabilities from LLMs by Leveraging Negative DataYiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan et al.AAAI 2024 · 31 citations
- Better Correlation and Robustness: A Distribution-Balanced Self-Supervised Learning Framework for Automatic Dialogue EvaluationPeiwen Yuan, Xinglin Wang, Jiayi Shi, Bin Sun et al.NeurIPS 2023 · 5 citations
- Focused Large Language Models are Stable Many-Shot LearnersPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang et al.EMNLP 2024
Builds on9
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng et al.AAAI 2020 · 354 citations
- Peer Collaborative Learning for Online Knowledge DistillationGuile Wu, Shaogang GongAAAI 2021 · 150 citations
- You Impress Me: Dialogue Generation via Mutual Persona PerceptionQian Liu, Yihong Chen, Bei Chen, Jian-Guang Lou et al.ACL 2020 · 144 citations
- Data-dependent Gaussian Prior Objective for Language GenerationZuchao Li, Rui Wang, Kehai Chen, Masao Utiyama et al.ICLR 2020 · 68 citations
- Posterior-GAN: Towards Informative and Coherent Response Generation with Posterior Generative Adversarial NetworkShaoxiong Feng, Hongshen Chen, Kan Li, Dawei YinAAAI 2020 · 26 citations
Related papers
- Multi-View Feature Representation for Dialogue Generation with Bidirectional DistillationShaoxiong Feng, Xuancheng Ren, Kan Li, Xu SunAAAI 2021 · 13 citations
- Online Knowledge Distillation via Collaborative LearningQiushan Guo, Xinjiang Wang, Yichao Wu, Zhipeng Yu et al.CVPR 2020
- CoLV: A Collaborative Latent Variable Model for Knowledge-Grounded Dialogue GenerationHaolan Zhan, Lei Shen, Hongshen Chen, Hainan ZhangEMNLP 2021 · 15 citations
- Weighted Mutual Learning with Diversity-Driven Model CompressionMiao Zhang, Li Wang, David Campos, Wei Huang et al.NeurIPS 2022 · 10 citations
- Dialogue Distillation: Open-Domain Dialogue Augmentation Using Unpaired DataRongsheng Zhang, Yinhe Zheng, Jianzhi Shao, Xiaoxi Mao et al.EMNLP 2020 · 25 citations
