Neural Machine Translation with Joint Representation
Yanyang Li, Qiang Wang, Tong Xiao, Tongran Liu, Jingbo Zhu
Abstract
Though early successes of Statistical Machine Translation (SMT) systems are attributed in part to the explicit modelling of the interaction between any two source and target units, e.g., alignment, the recent Neural Machine Translation (NMT) systems resort to the attention which partially encodes the interaction for efficiency. In this paper, we employ Joint Representation that fully accounts for each possible interaction. We sidestep the inefficiency issue by refining representations with the proposed efficient attention operation. The resulting Reformer models offer a new Sequence-to-Sequence modelling paradigm besides the Encoder-Decoder framework and outperform the Transformer baseline in either the small scale IWSLT14 German-English, English-German and IWSLT15 Vietnamese-English or the large scale NIST12 Chinese-English translation tasks by about 1 BLEU point. We also propose a systematic model scaling approach, allowing the Reformer model to beat the state-of-the-art Transformer in IWSLT14 German-English and NIST12 Chinese-English with about 50% fewer parameters. The code is publicly available at https://github.com/lyy1994/reformer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 111ec74c-4d63-4039-ad97-8ef085b8444aCited by top-tier papers1
Ask how each one uses itRelated papers
- Recurrent Attention for Neural Machine TranslationJiali Zeng, Shuangzhi Wu, Yongjing Yin, Yufan Jiang et al.EMNLP 2021
- Rephrasing the Reference for Non-autoregressive Machine TranslationChenze Shao, Jinchao Zhang, Jie Zhou, Yang FengAAAI 2023 · 6 citations
- Learning Source Phrase Representations for Neural Machine TranslationHongfei Xu, Josef van Genabith, Deyi Xiong, Qiuhui Liu et al.ACL 2020 · 18 citations
- ClusterFormer: Neural Clustering Attention for Efficient and Effective TransformerNingning Wang, Guobing Gan, Peng Zhang, Shuai Zhang et al.ACL 2022 · 24 citations
- Multiscale Collaborative Deep Models for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Yue Zhang et al.ACL 2020 · 27 citations
