Data Rejuvenation: Exploiting Inactive Training Examples for Neural Machine Translation
Wenxiang Jiao, Xing Wang, Shilin He, Irwin King, Michael R. Lyu, Zhaopeng Tu
摘要
Large-scale training datasets lie at the core of the recent success of neural machine translation (NMT) models. However, the complex patterns and potential noises in the large-scale data make training NMT models difficult. In this work, we explore to identify the inactive training examples which contribute less to the model performance, and show that the existence of inactive examples depends on the data distribution. We further introduce data rejuvenation to improve the training of NMT models on large-scale datasets by exploiting inactive examples. The proposed framework consists of three phases. First, we train an identification model on the original training data, and use it to distinguish inactive examples and active examples by their sentence-level output probabilities. Then, we train a rejuvenation model on the active examples, which is used to re-label the inactive examples with forwardtranslation. Finally, the rejuvenated examples and the active examples are combined to train the final NMT model. Experimental results on WMT14 English-German and English-French datasets show that the proposed data rejuvenation consistently and significantly improves performance for several strong NMT models. Extensive analyses reveal that our approach stabilizes and accelerates the training process of NMT models, resulting in final models with better generalization capability. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine TranslationWenxuan Wang, Wenxiang Jiao, Yongchang Hao, Xing Wang 等ACL 2022 · 被引用 32 次
- Towards Reliable Neural Machine Translation with Consistency-Aware Meta-LearningRongxiang Weng, Qiang Wang, Wensen Cheng, Changfeng Zhu 等AAAI 2023 · 被引用 3 次
- Self-Training Sampling with Monolingual Data Uncertainty for Neural Machine TranslationWenxiang Jiao, Xing Wang, Zhaopeng Tu, Shuming Shi 等ACL 2021
- Rejuvenating Low-Frequency Words: Making the Most of Parallel Data in Non-Autoregressive TranslationLiang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong 等ACL 2021
- Redistributing Low-Frequency Words: Making the Most of Monolingual Data in Non-Autoregressive TranslationLiang Ding, Longyue Wang, Shuming Shi, Dacheng Tao 等ACL 2022
它引用的顶会 Paper6
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 被引用 235 次
- Norm-Based Curriculum Learning for Neural Machine TranslationXuebo Liu, Houtim Lai, Derek F. Wong, Lidia S. ChaoACL 2020 · 被引用 97 次
- On the Inference Calibration of Neural Machine TranslationShuo Wang, Zhaopeng Tu, Shuming Shi, Yang LiuACL 2020 · 被引用 66 次
- Understanding Why Neural Networks Generalize Well Through GSNR of ParametersJinlong Liu, Yunzhi Bai, Guoqing Jiang, Ting Chen 等ICLR 2020 · 被引用 60 次
- Simplify-Then-Translate: Automatic Preprocessing for Black-Box TranslationSneha Mehta, Bahareh Azarnoush, Boris Chen, Avneesh Saluja 等AAAI 2020 · 被引用 19 次
相关 Paper
- Data Diversification: A Simple Strategy For Neural Machine TranslationXuan-Phi Nguyen, Shafiq R. Joty, Kui Wu, Ai Ti AwNeurIPS 2020 · 被引用 75 次
- Learning Source Phrase Representations for Neural Machine TranslationHongfei Xu, Josef van Genabith, Deyi Xiong, Qiuhui Liu 等ACL 2020 · 被引用 18 次
- Reinforced Curriculum Learning on Pre-Trained Neural Machine Translation ModelsMingjun Zhao, Haijiang Wu, Di Niu, Xiaoli WangAAAI 2020 · 被引用 46 次
- Meta Back-TranslationHieu Pham, Xinyi Wang, Yiming Yang, Graham NeubigICLR 2021 · 被引用 26 次
- Mirror-Generative Neural Machine TranslationZaixiang Zheng, Hao Zhou, Shujian Huang, Lei Li 等ICLR 2020 · 被引用 37 次
