Domain adapted machine translation: What does catastrophic forgetting forget and why?
Danielle Saunders, Steve DeNeefe
Abstract
Neural Machine Translation (NMT) models can be specialized by domain adaptation, often involving fine-tuning on a dataset of interest. This process risks catastrophic forgetting: rapid loss of generic translation quality. Forgetting has been widely observed, with many mitigation methods proposed. However, the causes of forgetting and the relationship between forgetting and adaptation data are under-explored. This paper takes a novel approach to understanding catastrophic forgetting during NMT adaptation by investigating the impact of the data. We provide a first investigation of what is forgotten, and why. We examine the relationship between forgetting and the in-domain data, and show that the amount and type of forgetting is linked to that data's target vocabulary coverage. Our findings pave the way toward better informed NMT domain adaptation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 984119c3-5446-4a1f-a97b-54bcccf30a65Builds on3
- Overcoming Catastrophic Forgetting beyond Continual Learning: Balanced Training for Neural Machine TranslationChenze Shao, Yang FengACL 2022 · 40 citations
- Unsupervised Domain Clusters in Pretrained Language ModelsRoee Aharoni, Yoav GoldbergACL 2020 · 13 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
Related papers
- MetaMT, a Meta Learning Method Leveraging Multiple Domain Data for Low Resource Machine TranslationRumeng Li, Xun Wang, Hong YuAAAI 2020 · 42 citations
- Continual Learning of Neural Machine Translation within Low Forgetting Risk RegionsShuhao Gu, Bojie Hu, Yang FengEMNLP 2022 · 11 citations
- Soft Alignment Objectives for Robust Adaptation of Language GenerationMichal Stefánik, Marek Kadlcík, Petr SojkaACL 2023 · 1 citation
- Distilling Multiple Domains for Neural Machine TranslationAnna Currey, Prashant Mathur, Georgiana DinuEMNLP 2020 · 19 citations
- Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine TranslationWenxuan Wang, Wenxiang Jiao, Yongchang Hao, Xing Wang et al.ACL 2022 · 32 citations
