Finding Sparse Structures for Domain Specific Neural Machine Translation
Jianze Liang, Chengqi Zhao, Mingxuan Wang, Xipeng Qiu, Lei Li
Abstract
Neural machine translation often adopts the fine-tuning approach to adapt to specific domains. However, nonrestricted fine-tuning can easily degrade on the general domain and over-fit to the target domain. To mitigate the issue, we propose PRUNE-TUNE, a novel domain adaptation method via gradual pruning. It learns tiny domain-specific sub-networks during fine-tuning on new domains. PRUNE-TUNE alleviates the over-fitting and the degradation problem without model modification. Furthermore, PRUNE-TUNE is able to sequentially learn a single network with multiple disjoint domainspecific sub-networks for multiple domains. Empirical experiment results show that PRUNE-TUNE outperforms several strong competitors in the target domain test set without sacrificing the quality on the general domain in both single and multi-domain settings. The source code and data are available at https://github.com/ohlionel/Prune-Tune . * The work was done while JL was an intern at ByteDance AI Lab.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- On the Effectiveness of Parameter-Efficient Fine-TuningZihao Fu, Haoran Yang, Anthony Man-Cho So, Wai Lam et al.AAAI 2023 · 234 citations
- Overcoming Catastrophic Forgetting beyond Continual Learning: Balanced Training for Neural Machine TranslationChenze Shao, Yang FengACL 2022 · 40 citations
- Improving the Cross-Lingual Generalisation in Visual Question AnsweringFarhad Nooralahzadeh, Rico SennrichAAAI 2023 · 8 citations
- Continual Learning with Semi-supervised Contrastive Distillation for Incremental Neural Machine TranslationYunlong Liang, Fandong Meng, Jiaan Wang, Jinan Xu et al.ACL 2024 · 7 citations
- Continual Knowledge Distillation for Neural Machine TranslationYuanchi Zhang, Peng Li, Maosong Sun, Yang LiuACL 2023 · 5 citations
Builds on4
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsCheolhyoung Lee, Kyunghyun Cho, Wanmo KangICLR 2020 · 233 citations
- Revisiting Few-sample BERT Fine-tuningTianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger et al.ICLR 2021 · 172 citations
- Learning Sparse Sharing Architectures for Multiple TasksTianxiang Sun, Yunfan Shao, Xiaonan Li, Pengfei Liu et al.AAAI 2020 · 155 citations
- Learning a Multi-Domain Curriculum for Neural Machine TranslationWei Wang, Ye Tian, Jiquan Ngiam, Yinfei Yang et al.ACL 2020 · 32 citations
Related papers
- Gradient-based Gradual Pruning for Language-Specific Multilingual Neural Machine TranslationDan He, Minh-Quang Pham, Thanh-Le Ha, Marco TurchiEMNLP 2023 · 2 citations
- MetaMT, a Meta Learning Method Leveraging Multiple Domain Data for Low Resource Machine TranslationRumeng Li, Xun Wang, Hong YuAAAI 2020 · 42 citations
- Continual Learning for Multilingual Neural Machine Translation via Dual Importance-based Model DivisionJunpeng Liu, Kaiyu Huang, Hao Yu, Jiuyi Li et al.EMNLP 2023 · 4 citations
- Distilling Multiple Domains for Neural Machine TranslationAnna Currey, Prashant Mathur, Georgiana DinuEMNLP 2020 · 19 citations
- Knowledge Transfer in Incremental Learning for Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Jin Ma, Ting Yao et al.ACL 2023 · 17 citations
