Self-Training Sampling with Monolingual Data Uncertainty for Neural Machine Translation
Wenxiang Jiao, Xing Wang, Zhaopeng Tu, Shuming Shi, Michael R. Lyu, Irwin King
摘要
Self-training has proven effective for improving NMT performance by augmenting model training with synthetic parallel data. The common practice is to construct synthetic data based on a randomly sampled subset of large-scale monolingual data, which we empirically show is sub-optimal. In this work, we propose to improve the sampling procedure by selecting the most informative monolingual sentences to complement the parallel data. To this end, we compute the uncertainty of monolingual sentences using the bilingual dictionary extracted from the parallel data. Intuitively, monolingual sentences with lower uncertainty generally correspond to easy-to-translate patterns which may not provide additional gains. Accordingly, we design an uncertainty-based sampling strategy to efficiently exploit the monolingual data for self-training, in which monolingual sentences with higher uncertainty would be sampled with higher probability. Experimental results on large-scale WMT English⇒German and English⇒Chinese datasets demonstrate the effectiveness of the proposed approach. Extensive analyses suggest that emphasizing the learning on uncertain monolingual sentences by our approach does improve the translation quality of high-uncertainty sentences and also benefits the prediction of low-frequency words at the target side. 1 * Work was mainly done when Wenxiang Jiao was interning at Tencent AI Lab.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine TranslationWenxuan Wang, Wenxiang Jiao, Yongchang Hao, Xing Wang 等ACL 2022 · 被引用 32 次
- Teacher Forcing Recovers Reward Functions for Text GenerationYongchang Hao, Yuxin Liu, Lili MouNeurIPS 2022 · 被引用 24 次
- Improving Simultaneous Machine Translation with Monolingual DataHexuan Deng, Liang Ding, Xuebo Liu, Meishan Zhang 等AAAI 2023 · 被引用 19 次
- Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code ModelsShuzheng Gao, Wenxin Mao, Cuiyun Gao, Li Li 等ICSE 2024 · 被引用 15 次
- DuNST: Dual Noisy Self Training for Semi-Supervised Controllable Text GenerationYuxi Feng, Xiaoyuan Yi, Xiting Wang, Laks V. S. Lakshmanan 等ACL 2023 · 被引用 2 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Revisiting Self-Training for Neural Sequence GenerationJunxian He, Jiatao Gu, Jiajun Shen, Marc'Aurelio RanzatoICLR 2020 · 被引用 294 次
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 被引用 235 次
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 被引用 182 次
- Uncertainty-Aware Curriculum Learning for Neural Machine TranslationYikai Zhou, Baosong Yang, Derek F. Wong, Yu Wan 等ACL 2020 · 被引用 78 次
相关 Paper
- Uncertainty-Aware Semantic Augmentation for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Rongxiang Weng 等EMNLP 2020 · 被引用 20 次
- Uncertainty-Aware Balancing for Multilingual and Multi-Domain Neural Machine Translation TrainingMinghao Wu, Yitong Li, Meng Zhang, Liangyou Li 等EMNLP 2021 · 被引用 10 次
- Meta Back-TranslationHieu Pham, Xinyi Wang, Yiming Yang, Graham NeubigICLR 2021 · 被引用 26 次
- Dynamic Data Selection and Weighting for Iterative Back-TranslationZi-Yi Dou, Antonios Anastasopoulos, Graham NeubigEMNLP 2020 · 被引用 46 次
- Bridging the Data Gap between Training and Inference for Unsupervised Neural Machine TranslationZhiwei He, Xing Wang, Rui Wang, Shuming Shi 等ACL 2022
