Redistributing Low-Frequency Words: Making the Most of Monolingual Data in Non-Autoregressive Translation
Liang Ding, Longyue Wang, Shuming Shi, Dacheng Tao, Zhaopeng Tu
摘要
Knowledge distillation (KD) is the preliminary step for training non-autoregressive translation (NAT) models, which eases the training of NAT models at the cost of losing important information for translating low-frequency words. In this work, we provide an appealing alternative for NAT – monolingual KD, which trains NAT student on external monolingual data with AT teacher trained on the original bilingual data. Monolingual KD is able to transfer both the knowledge of the original bilingual data (implicitly encoded in the trained AT teacher model) and that of the new monolingual data to the NAT student model. Extensive experiments on eight WMT benchmarks over two advanced NAT models show that monolingual KD consistently outperforms the standard KD by improving low-frequency word translation, without introducing any computational cost. Monolingual KD enjoys desirable expandability, which can be further enhanced (when given more computational budget) by combining with the standard KD, a reverse monolingual KD, or enlarging the scale of monolingual data. Extensive analyses demonstrate that these techniques can be used together profitably to further recall the useful information lost in the standard KD. Encouragingly, combining with standard KD, our approach achieves 30.4 and 34.1 BLEU points on the WMT14 English-German and German-English datasets, respectively. Our code and trained models are freely available at https://github.com/alphadl/RLFW-NAT.mono.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Improving Simultaneous Machine Translation with Monolingual DataHexuan Deng, Liang Ding, Xuebo Liu, Meishan Zhang 等AAAI 2023 · 被引用 19 次
- Multi-Step Denoising Scheduled Sampling: Towards Alleviating Exposure Bias for Diffusion ModelsZhiyao Ren, Yibing Zhan, Liang Ding, Gaoang Wang 等AAAI 2024 · 被引用 15 次
- Divide, Conquer, and Combine: Mixture of Semantic-Independent Experts for Zero-Shot Dialogue State TrackingQingyue Wang, Liang Ding, Yanan Cao, Yibing Zhan 等ACL 2023 · 被引用 5 次
- XLM-D: Decorate Cross-lingual Pre-training Model as Non-Autoregressive Neural Machine TranslationYong Wang, Shilin He, Guanhua Chen, Yun Chen 等EMNLP 2022 · 被引用 4 次
- Zero-shot Sharpness-Aware Quantization for Pre-trained Language ModelsMiaoxi Zhu, Qihuang Zhong, Li Shen, Liang Ding 等EMNLP 2023 · 被引用 2 次
它引用的顶会 Paper11
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 被引用 235 次
- Aligned Cross Entropy for Non-Autoregressive Machine TranslationMarjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, Omer LevyICML 2020 · 被引用 121 次
- Minimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine TranslationChenze Shao, Jinchao Zhang, Yang Feng, Fandong Meng 等AAAI 2020 · 被引用 95 次
- Order-Agnostic Cross Entropy for Non-Autoregressive Machine TranslationCunxiao Du, Zhaopeng Tu, Jing JiangICML 2021 · 被引用 93 次
- Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine TranslationJunliang Guo, Xu Tan, Linli Xu, Tao Qin 等AAAI 2020 · 被引用 91 次
相关 Paper
- Rejuvenating Low-Frequency Words: Making the Most of Parallel Data in Non-Autoregressive TranslationLiang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong 等ACL 2021
- Understanding and Improving Lexical Choice in Non-Autoregressive TranslationLiang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong 等ICLR 2021 · 被引用 44 次
- Towards Understanding and Improving Knowledge Distillation for Neural Machine TranslationSongming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen 等ACL 2023 · 被引用 8 次
- Data Diversification: A Simple Strategy For Neural Machine TranslationXuan-Phi Nguyen, Shafiq R. Joty, Kui Wu, Ai Ti AwNeurIPS 2020 · 被引用 75 次
- Revisiting Knowledge Distillation for Autoregressive Language ModelsQihuang Zhong, Liang Ding, Li Shen, Juhua Liu 等ACL 2024
