Multiscale Collaborative Deep Models for Neural Machine Translation
Xiangpeng Wei, Heng Yu, Yue Hu, Yue Zhang, Rongxiang Weng, Weihua Luo
摘要
Recent evidence reveals that Neural Machine Translation (NMT) models with deeper neural networks can be more effective but are difficult to train. In this paper, we present a MultiScale Collaborative (MSC) framework to ease the training of NMT models that are substantially deeper than those used previously. We explicitly boost the gradient back-propagation from top to bottom levels by introducing a block-scale collaboration mechanism into deep NMT models. Then, instead of forcing the whole encoder stack directly learns a desired representation, we let each encoder block learns a fine-grained representation and enhance it by encoding spatial dependencies using a context-scale collaboration. We provide empirical evidence showing that the MSC nets are easy to optimize and can obtain improvements of translation quality from considerably increased depth. On IWSLT translation tasks with three translation directions, our extremely deep models (with 72-layer encoders) surpass strong baselines by +2.2 +3.1 BLEU points. In addition, our deep MSC achieves a BLEU score of 30.56 on WMT14 English-to-German task that significantly outperforms state-of-the-art deep NMT models. We have included the source code in supplementary materials.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- On Learning Universal Representations Across LanguagesXiangpeng Wei, Rongxiang Weng, Yue Hu, Luxi Xing 等ICLR 2021 · 被引用 93 次
- Learning Light-Weight Translation Models from Deep TransformerBei Li, Ziyang Wang, Hui Liu, Quan Du 等AAAI 2021 · 被引用 44 次
- ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence GenerationBei Li, Quan Du, Tao Zhou, Yi Jing 等ACL 2022 · 被引用 43 次
- Shallow-to-Deep Training for Neural Machine TranslationBei Li, Ziyang Wang, Hui Liu, Yufan Jiang 等EMNLP 2020 · 被引用 41 次
- Deep Transformers with Latent DepthXian Li, Asa Cooper Stickland, Yuqing Tang, Xiang KongNeurIPS 2020 · 被引用 32 次
相关 Paper
- Jointly Masked Sequence-to-Sequence Model for Non-Autoregressive Neural Machine TranslationJunliang Guo, Linli Xu, Enhong ChenACL 2020 · 被引用 56 次
- Data Diversification: A Simple Strategy For Neural Machine TranslationXuan-Phi Nguyen, Shafiq R. Joty, Kui Wu, Ai Ti AwNeurIPS 2020 · 被引用 75 次
- BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine TranslationHaoran Xu, Benjamin Van Durme, Kenton W. MurrayEMNLP 2021 · 被引用 55 次
- Confidence Based Bidirectional Global Context Aware Training Framework for Neural Machine TranslationChulun Zhou, Fandong Meng, Jie Zhou, Min Zhang 等ACL 2022 · 被引用 20 次
- Norm-Based Curriculum Learning for Neural Machine TranslationXuebo Liu, Houtim Lai, Derek F. Wong, Lidia S. ChaoACL 2020 · 被引用 97 次
