Domain adapted machine translation: What does catastrophic forgetting forget and why?
Danielle Saunders, Steve DeNeefe
摘要
Neural Machine Translation (NMT) models can be specialized by domain adaptation, often involving fine-tuning on a dataset of interest. This process risks catastrophic forgetting: rapid loss of generic translation quality. Forgetting has been widely observed, with many mitigation methods proposed. However, the causes of forgetting and the relationship between forgetting and adaptation data are under-explored. This paper takes a novel approach to understanding catastrophic forgetting during NMT adaptation by investigating the impact of the data. We provide a first investigation of what is forgotten, and why. We examine the relationship between forgetting and the in-domain data, and show that the amount and type of forgetting is linked to that data's target vocabulary coverage. Our findings pave the way toward better informed NMT domain adaptation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Overcoming Catastrophic Forgetting beyond Continual Learning: Balanced Training for Neural Machine TranslationChenze Shao, Yang FengACL 2022 · 被引用 40 次
- Unsupervised Domain Clusters in Pretrained Language ModelsRoee Aharoni, Yoav GoldbergACL 2020 · 被引用 13 次
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
相关 Paper
- MetaMT, a Meta Learning Method Leveraging Multiple Domain Data for Low Resource Machine TranslationRumeng Li, Xun Wang, Hong YuAAAI 2020 · 被引用 42 次
- Continual Learning of Neural Machine Translation within Low Forgetting Risk RegionsShuhao Gu, Bojie Hu, Yang FengEMNLP 2022 · 被引用 11 次
- Soft Alignment Objectives for Robust Adaptation of Language GenerationMichal Stefánik, Marek Kadlcík, Petr SojkaACL 2023 · 被引用 1 次
- Distilling Multiple Domains for Neural Machine TranslationAnna Currey, Prashant Mathur, Georgiana DinuEMNLP 2020 · 被引用 19 次
- Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine TranslationWenxuan Wang, Wenxiang Jiao, Yongchang Hao, Xing Wang 等ACL 2022 · 被引用 32 次
