Prediction Difference Regularization against Perturbation for Neural Machine Translation
Dengji Guo, Zhengrui Ma, Min Zhang, Yang Feng
Abstract
Regularization methods applying input perturbation have drawn considerable attention and have been frequently explored for NMT tasks in recent years. Despite their simplicity and effectiveness, we argue that these methods are limited by the under-fitting of training data. In this paper, we utilize prediction difference for ground-truth tokens to analyze the fitting of token-level samples and find that underfitting is almost as common as over-fitting. We introduce prediction difference regularization (PD-R), a simple and effective method that can reduce over-fitting and under-fitting at the same time. For all token-level samples, PD-R minimizes the prediction difference between the original pass and the input-perturbed pass, making the model less sensitive to small input changes, thus more robust to both perturbations and under-fitted training data. Experiments on three widely used WMT translation tasks show that our approach can significantly improve over existing perturbation regularization methods. On WMT16 En-De task, our model achieves 1.80 SacreBLEU improvement over vanilla transformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine TranslationZhaocong Li, Xuebo Liu, Derek F. Wong, Lidia S. Chao et al.EMNLP 2022 · 20 citations
- Understanding and Bridging the Modality Gap for Speech TranslationQingkai Fang, Yang FengACL 2023 · 12 citations
- STEMM: Self-learning with Speech-text Manifold Mixup for Speech TranslationQingkai Fang, Rong Ye, Lei Li, Yang Feng et al.ACL 2022
Builds on4
- Evaluating the Robustness of Neural Language Models to Input PerturbationsMilad Moradi, Matthias SamwaldEMNLP 2021 · 64 citations
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 17 citations
- STEMM: Self-learning with Speech-text Manifold Mixup for Speech TranslationQingkai Fang, Rong Ye, Lei Li, Yang Feng et al.ACL 2022
- Beyond Noise: Mitigating the Impact of Fine-grained Semantic Divergences on Neural Machine TranslationEleftheria Briakou, Marine CarpuatACL 2021
Related papers
- Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine TranslationWenxuan Wang, Wenxiang Jiao, Yongchang Hao, Xing Wang et al.ACL 2022 · 32 citations
- TLM: Token-Level Masking for TransformersYangjun Wu, Kebin Fang, Dongxiang Zhang, Han Wang et al.EMNLP 2023 · 2 citations
- Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation ModelsTianjian Li, Haoran Xu, Philipp Koehn, Daniel Khashabi et al.ICLR 2024 · 6 citations
- Data Diversification: A Simple Strategy For Neural Machine TranslationXuan-Phi Nguyen, Shafiq R. Joty, Kui Wu, Ai Ti AwNeurIPS 2020 · 75 citations
- Token-level Adaptive Training for Neural Machine TranslationShuhao Gu, Jinchao Zhang, Fandong Meng, Yang Feng et al.EMNLP 2020 · 32 citations
