Learning Source Phrase Representations for Neural Machine Translation
Hongfei Xu, Josef van Genabith, Deyi Xiong, Qiuhui Liu, Jingyi Zhang
Abstract
The Transformer translation model (Vaswani et al., 2017) based on a multi-head attention mechanism can be computed effectively in parallel and has significantly pushed forward the performance of Neural Machine Translation (NMT). Though intuitively the attentional network can connect distant words via shorter network paths than RNNs, empirical analysis demonstrates that it still has difficulty in fully capturing long-distance dependencies (Tang et al., 2018) . Considering that modeling phrases instead of words has significantly improved the Statistical Machine Translation (SMT) approach through the use of larger translation blocks ("phrases") and its reordering ability, modeling NMT at phrase level is an intuitive proposal to help the model capture long-distance relationships. In this paper, we first propose an attentive phrase representation generation mechanism which is able to generate phrase representations from corresponding token representations. In addition, we incorporate the generated phrase representations into the Transformer translation model to enhance its ability to capture long-distance relationships. In our experiments, we obtain significant improvements on the WMT 14 English-German and English-French tasks on top of the strong Transformer baseline, which shows the effectiveness of our approach. Our approach helps Transformer Base models perform at the level of Transformer Big models, and even significantly better for long sentences, but with substantially fewer parameters and training steps. The fact that phrase representations help even in the big setting further supports our conjecture that they make a valuable contribution to long-distance relations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a5b383b-6cfc-4b68-a175-a6d66d85057eCited by top-tier papers3
- Cross2StrA: Unpaired Cross-lingual Image Captioning with Cross-lingual Cross-modal Structure-pivoted AlignmentShengqiong Wu, Hao Fei, Wei Ji, Tat-Seng ChuaACL 2023 · 42 citations
- CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based VerificationYuchen Tian, Weixiang Yan, Qian Yang, Xuandong Zhao et al.AAAI 2025 · 41 citations
- Multi-Head Highly Parallelized LSTM Decoder for Neural Machine TranslationHongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong et al.ACL 2021
Related papers
- Multi-Unit Transformers for Neural Machine TranslationJianhao Yan, Fandong Meng, Jie ZhouEMNLP 2020 · 21 citations
- Explicit Sentence Compression for Neural Machine TranslationZuchao Li, Rui Wang, Kehai Chen, Masao Utiyama et al.AAAI 2020 · 31 citations
- Generating Diverse Translation by Manipulating Multi-Head AttentionZewei Sun, Shujian Huang, Hao-Ran Wei, Xinyu Dai et al.AAAI 2020 · 36 citations
- Learning Multiscale Transformer Models for Sequence GenerationBei Li, Tong Zheng, Yi Jing, Chengbo Jiao et al.ICML 2022 · 15 citations
- An Efficient Transformer Decoder with Compressed Sub-layersYanyang Li, Ye Lin, Tong Xiao, Jingbo ZhuAAAI 2021 · 32 citations
