Multilingual Machine Translation with Hyper-Adapters
Christos Baziotis, Mikel Artetxe, James Cross, Shruti Bhosale
Abstract
Multilingual machine translation suffers from negative interference across languages. A common solution is to relax parameter sharing with language-specific modules like adapters. However, adapters of related languages are unable to transfer information, and their total number of parameters becomes prohibitively expensive as the number of languages grows. In this work, we overcome these drawbacks using hyper-adapters – hyper-networks that generate adapters from language and layer embeddings. While past work had poor results when scaling hyper-networks, we propose a rescaling fix that significantly improves convergence and enables training larger hyper-networks. We find that hyper-adapters are more parameter efficient than regular adapters, reaching the same performance with up to 12 times less parameters. When using the same number of parameters and FLOPS, our approach consistently outperforms regular adapters. Also, hyper-adapters converge faster than alternative approaches and scale better than regular dense networks. Our analysis shows that hyper-adapters learn to encode language relatedness, enabling positive transfer across languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cbfe3fa8-1e29-4560-8c30-05007418f6dbCited by top-tier papers9
- EBBS: An Ensemble with Bi-Level Beam Search for Zero-Shot Machine TranslationYuqiao Wen, Behzad Shayegh, Chenyang Huang, Yanshuai Cao et al.AAAI 2025 · 8 citations
- Learning Language-Specific Layers for Multilingual Machine TranslationTelmo Pires, Robin M. Schmidt, Yi-Hsiu Liao, Stephan PeitzACL 2023 · 8 citations
- From Instance Training to Instruction Learning: Task Adapters Generation from InstructionsHuanxuan Liao, Shizhu He, Yao Xu, Yuanzhe Zhang et al.NeurIPS 2024 · 6 citations
- Condensing Multilingual Knowledge with Lightweight Language-Specific ModulesHaoran Xu, Weiting Tan, Shuyue Stella Li, Yunmo Chen et al.EMNLP 2023 · 3 citations
- Gradient-based Gradual Pruning for Language-Specific Multilingual Neural Machine TranslationDan He, Minh-Quang Pham, Thanh-Le Ha, Marco TurchiEMNLP 2023 · 2 citations
Builds on9
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 213 citations
- Principled Weight Initialization for HypernetworksOscar Chang, Lampros Flokas, Hod LipsonICLR 2020 · 87 citations
- HyperGrid Transformers: Towards A Single Model for Multiple TasksYi Tay, Zhe Zhao, Dara Bahri, Donald Metzler et al.ICLR 2021 · 44 citations
- MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual TransferJonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian RuderEMNLP 2020 · 36 citations
- VL-ADAPTER: Parameter-Efficient Transfer Learning for Vision-and-Language TasksYi-Lin Sung, Jaemin Cho, Mohit BansalCVPR 2022 · 22 citations
Related papers
- Language Constrained Multimodal Hyper Adapter For Many-to-Many Multimodal SummarizationNayu Liu, Fanglong Yao, Haoran Luo, Yong Yang et al.ACL 2025
- Hyper-X: A Unified Hypernetwork for Multi-Task Multilingual TransferAhmet Üstün, Arianna Bisazza, Gosse Bouma, Gertjan van Noord et al.EMNLP 2022 · 13 citations
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of MultilingualityShayne Longpre, Sneha Kudugunta, Niklas Muennighoff, I-Hung Hsu et al.ICLR 2026 · 20 citations
- Beyond Shared Vocabulary: Increasing Representational Word Similarities across Languages for Multilingual Machine TranslationDi Wu, Christof MonzEMNLP 2023 · 5 citations
- Parameter-efficient Multi-task Fine-tuning for Transformers via Shared HypernetworksRabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, James HendersonACL 2021
