Norm-Based Curriculum Learning for Neural Machine Translation
Xuebo Liu, Houtim Lai, Derek F. Wong, Lidia S. Chao
摘要
A neural machine translation (NMT) system is expensive to train, especially with highresource settings. As the NMT architectures become deeper and wider, this issue gets worse and worse. In this paper, we aim to improve the efficiency of training an NMT by introducing a novel norm-based curriculum learning method. We use the norm (aka length or module) of a word embedding as a measure of 1) the difficulty of the sentence, 2) the competence of the model, and 3) the weight of the sentence. The normbased sentence difficulty takes the advantages of both linguistically motivated and modelbased sentence difficulties. It is easy to determine and contains learning-dependent features. The norm-based model competence makes NMT learn the curriculum in a fully automated way, while the norm-based sentence weight further enhances the learning of the vector representation of the NMT. Experimental results for the WMT'14 English-German and WMT'17 Chinese-English translation tasks demonstrate that the proposed method outperforms strong baselines in terms of BLEU score (+1.17/+1.56) and training speedup (2.22x/3.33x).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Uncertainty-Aware Curriculum Learning for Neural Machine TranslationYikai Zhou, Baosong Yang, Derek F. Wong, Yu Wan 等ACL 2020 · 被引用 78 次
- Spatio-Temporal Trajectory Similarity Learning in Road NetworksZiquan Fang, Yuntao Du, Xinjun Zhu, Danlei Hu 等KDD 2022 · 被引用 68 次
- Hybrid Curriculum Learning for Emotion Recognition in ConversationLin Yang, Yi Shen, Yue Mao, Longjun CaiAAAI 2022 · 被引用 64 次
- Meta-Curriculum Learning for Domain Adaptation in Neural Machine TranslationRunzhe Zhan, Xuebo Liu, Derek F. Wong, Lidia S. ChaoAAAI 2021 · 被引用 50 次
- M2DF: Multi-grained Multi-curriculum Denoising Framework for Multimodal Aspect-based Sentiment AnalysisFei Zhao, Chunhui Li, Zhen Wu, Yawen Ouyang 等EMNLP 2023 · 被引用 42 次
它引用的顶会 Paper1
相关 Paper
- Efficient Pre-training of Masked Language Model via Concept-based Curriculum MaskingMingyu Lee, Jun-Hyung Park, Junho Kim, Kang-Min Kim 等EMNLP 2022 · 被引用 8 次
- Multiscale Collaborative Deep Models for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Yue Zhang 等ACL 2020 · 被引用 27 次
- Fast and Accurate Neural Machine Translation with Translation MemoryQiuxiang He, Guoping Huang, Qu Cui, Li Li 等ACL 2021
- Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine TranslationJunliang Guo, Xu Tan, Linli Xu, Tao Qin 等AAAI 2020 · 被引用 91 次
- Accelerating Neural Machine Translation with Partial Word Embedding CompressionFan Zhang, Mei Tu, Jinyao YanAAAI 2021 · 被引用 3 次
