Cross-Attention is All You Need: Adapting Pretrained Transformers for Machine Translation
Mozhdeh Gheini, Xiang Ren, Jonathan May
Abstract
We study the power of cross-attention in the Transformer architecture within the context of transfer learning for machine translation, and extend the findings of studies into crossattention when training from scratch. We conduct a series of experiments through finetuning a translation model on data where either the source or target language has changed. These experiments reveal that fine-tuning only the cross-attention parameters is nearly as effective as fine-tuning all parameters (i.e., the entire translation model). We provide insights into why this is the case and observe that limiting fine-tuning in this manner yields crosslingually aligned embeddings. The implications of this finding for researchers and practitioners include a mitigation of catastrophic forgetting, the potential for zero-shot translation, and the ability to extend machine translation models to several new language pairs with reduced parameter storage overhead. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5b972c6-001a-4270-add9-ef84eccf4f4fCited by top-tier papers19
- LIFT: Language-Interfaced Fine-Tuning for Non-language Machine Learning TasksTuan Dinh, Yuchen Zeng, Ruisu Zhang, Ziqian Lin et al.NeurIPS 2022 · 222 citations
- Parameter-Efficient Orthogonal Finetuning via Butterfly FactorizationWeiyang Liu, Zeju Qiu, Yao Feng, Yuliang Xiu et al.ICLR 2024 · 111 citations
- Causal Intervention for Human Trajectory Prediction with Cross Attention MechanismChunjiang Ge, Shiji Song, Gao HuangAAAI 2023 · 30 citations
- PLoP: Precise LoRA Placement for Efficient Finetuning of Large ModelsSoufiane Hayou, Nikhil Ghosh, Bin YuICLR 2026 · 14 citations
- Dynamic Layer Tying for Parameter-Efficient TransformersTamir David Hay, Lior WolfICLR 2024 · 13 citations
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 94 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
Related papers
- In Neural Machine Translation, What Does Transfer Learning Transfer?Alham Fikri Aji, Nikolay Bogoychev, Kenneth Heafield, Rico SennrichACL 2020 · 56 citations
- Translation Artifacts in Cross-lingual Transfer LearningMikel Artetxe, Gorka Labaka, Eneko AgirreEMNLP 2020 · 68 citations
- Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention ModelsChungpa Lee, Jy-yong Sohn, Kangwook LeeICML 2026 · 1 citation
- Hard-Coded Gaussian Attention for Neural Machine TranslationWeiqiu You, Simeng Sun, Mohit IyyerACL 2020 · 55 citations
- Composable Sparse Fine-Tuning for Cross-Lingual TransferAlan Ansell, Edoardo Maria Ponti, Anna Korhonen, Ivan VulicACL 2022
