Self-Training Sampling with Monolingual Data Uncertainty for Neural Machine Translation
Wenxiang Jiao, Xing Wang, Zhaopeng Tu, Shuming Shi, Michael R. Lyu, Irwin King
Abstract
Self-training has proven effective for improving NMT performance by augmenting model training with synthetic parallel data. The common practice is to construct synthetic data based on a randomly sampled subset of large-scale monolingual data, which we empirically show is sub-optimal. In this work, we propose to improve the sampling procedure by selecting the most informative monolingual sentences to complement the parallel data. To this end, we compute the uncertainty of monolingual sentences using the bilingual dictionary extracted from the parallel data. Intuitively, monolingual sentences with lower uncertainty generally correspond to easy-to-translate patterns which may not provide additional gains. Accordingly, we design an uncertainty-based sampling strategy to efficiently exploit the monolingual data for self-training, in which monolingual sentences with higher uncertainty would be sampled with higher probability. Experimental results on large-scale WMT English⇒German and English⇒Chinese datasets demonstrate the effectiveness of the proposed approach. Extensive analyses suggest that emphasizing the learning on uncertain monolingual sentences by our approach does improve the translation quality of high-uncertainty sentences and also benefits the prediction of low-frequency words at the target side. 1 * Work was mainly done when Wenxiang Jiao was interning at Tencent AI Lab.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine TranslationWenxuan Wang, Wenxiang Jiao, Yongchang Hao, Xing Wang et al.ACL 2022 · 32 citations
- Teacher Forcing Recovers Reward Functions for Text GenerationYongchang Hao, Yuxin Liu, Lili MouNeurIPS 2022 · 24 citations
- Improving Simultaneous Machine Translation with Monolingual DataHexuan Deng, Liang Ding, Xuebo Liu, Meishan Zhang et al.AAAI 2023 · 19 citations
- Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code ModelsShuzheng Gao, Wenxin Mao, Cuiyun Gao, Li Li et al.ICSE 2024 · 15 citations
- DuNST: Dual Noisy Self Training for Semi-Supervised Controllable Text GenerationYuxi Feng, Xiaoyuan Yi, Xiting Wang, Laks V. S. Lakshmanan et al.ACL 2023 · 2 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Revisiting Self-Training for Neural Sequence GenerationJunxian He, Jiatao Gu, Jiajun Shen, Marc'Aurelio RanzatoICLR 2020 · 294 citations
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 235 citations
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 182 citations
- Uncertainty-Aware Curriculum Learning for Neural Machine TranslationYikai Zhou, Baosong Yang, Derek F. Wong, Yu Wan et al.ACL 2020 · 78 citations
Related papers
- Uncertainty-Aware Semantic Augmentation for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Rongxiang Weng et al.EMNLP 2020 · 20 citations
- Uncertainty-Aware Balancing for Multilingual and Multi-Domain Neural Machine Translation TrainingMinghao Wu, Yitong Li, Meng Zhang, Liangyou Li et al.EMNLP 2021 · 10 citations
- Meta Back-TranslationHieu Pham, Xinyi Wang, Yiming Yang, Graham NeubigICLR 2021 · 26 citations
- Dynamic Data Selection and Weighting for Iterative Back-TranslationZi-Yi Dou, Antonios Anastasopoulos, Graham NeubigEMNLP 2020 · 46 citations
- Bridging the Data Gap between Training and Inference for Unsupervised Neural Machine TranslationZhiwei He, Xing Wang, Rui Wang, Shuming Shi et al.ACL 2022
