Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine Translation
Junliang Guo, Xu Tan, Linli Xu, Tao Qin, Enhong Chen, Tie-Yan Liu
摘要
Non-autoregressive translation (NAT) models remove the dependence on previous target tokens and generate all target tokens in parallel, resulting in significant inference speedup but at the cost of inferior translation accuracy compared to autoregressive translation (AT) models. Considering that AT models have higher accuracy and are easier to train than NAT models, and both of them share the same model configurations, a natural idea to improve the accuracy of NAT models is to transfer a well-trained AT model to an NAT model through fine-tuning. However, since AT and NAT models differ greatly in training strategy, straightforward fine-tuning does not work well. In this work, we introduce curriculum learning into fine-tuning for NAT. Specifically, we design a curriculum in the fine-tuning process to progressively switch the training from autoregressive generation to non-autoregressive generation. Experiments on four benchmark translation datasets show that the proposed method achieves good improvement (more than 1 BLEU score) over previous NAT baselines in terms of translation accuracy, and greatly speed up (more than 10 times) the inference process over AT baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Diffusion Language Models Are Versatile Protein LearnersXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue 等ICML 2024 · 被引用 113 次
- FastCorrect: Fast Error Correction with Edit Alignment for Automatic Speech RecognitionYichong Leng, Xu Tan, Linchen Zhu, Jin Xu 等NeurIPS 2021 · 被引用 84 次
- Guiding Non-Autoregressive Neural Machine Translation Decoding with Reordering InformationQiu Ran, Yankai Lin, Peng Li, Jie ZhouAAAI 2021 · 被引用 82 次
- Incorporating BERT into Parallel Sequence Decoding with AdaptersJunliang Guo, Zhirui Zhang, Linli Xu, Hao-Ran Wei 等NeurIPS 2020 · 被引用 72 次
- A Study of Non-autoregressive Model for Sequence GenerationYi Ren, Jinglin Liu, Xu Tan, Zhou Zhao 等ACL 2020 · 被引用 58 次
相关 Paper
- NAT4AT: Using Non-Autoregressive Translation Makes Autoregressive Translation Faster and BetterHuanran Zheng, Wei Zhu, Xiaoling WangWWW 2024 · 被引用 13 次
- Curriculum Learning for Biological Sequence Prediction: The Case of De Novo Peptide SequencingXiang Zhang, Jiaqi Wei, Zijie Qiu, Sheng Xu 等ICML 2025
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 被引用 235 次
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross 等ICLR 2021 · 被引用 154 次
- Learning to Rewrite for Non-Autoregressive Neural Machine TranslationXinwei Geng, Xiaocheng Feng, Bing QinEMNLP 2021 · 被引用 30 次
