Neuron Specialization: Leveraging Intrinsic Task Modularity for Multilingual Machine Translation
Shaomu Tan, Di Wu, Christof Monz
摘要
Training a unified multilingual model promotes knowledge transfer but inevitably introduces negative interference. Language-specific modeling methods show promise in reducing interference. However, they often rely on heuristics to distribute capacity and struggle to foster cross-lingual transfer via isolated modules. In this paper, we explore intrinsic task modularity within multilingual networks and leverage these observations to circumvent interference under multilingual translation. We show that neurons in the feed-forward layers tend to be activated in a language-specific manner. Meanwhile, these specialized neurons exhibit structural overlaps that reflect language proximity, which progress across layers. Based on these findings, we propose Neuron Specialization, an approach that identifies specialized neurons to modularize feed-forward layers and then continuously updates them through sparse networks. Extensive experiments show that our approach achieves consistent performance gains over strong baselines with additional analyses demonstrating reduced interference and increased knowledge transfer. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Style-Specific Neurons for Steering LLMs in Text Style TransferWen Lai, Viktor Hangya, Alexander FraserEMNLP 2024 · 被引用 5 次
- How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons PerspectiveShimao Zhang, Zhejian Lai, Xiang Liu, Shuaijie She 等AAAI 2026 · 被引用 4 次
- Neuron Empirical Gradient: Discovering and Quantifying Neurons' Global Linear ControllabilityXin Zhao, Zehui Jiang, Naoki YoshinagaACL 2025 · 被引用 3 次
- Linguistic Neuron Overlap Patterns to Facilitate Cross-lingual Transfer on Low-resource LanguagesYuemei Xu, Kexin Xu, Jian Zhou, Ling Hu 等EMNLP 2025 · 被引用 1 次
- MLAS-LoRA: Language-Aware Parameters Detection and LoRA-Based Knowledge Transfer for Multilingual Machine TranslationTianyu Dong, Bo Li, Jinsong Liu, Shaolin Zhu 等ACL 2025
它引用的顶会 Paper15
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual ModelsZirui Wang, Yulia Tsvetkov, Orhan Firat, Yuan CaoICLR 2021 · 被引用 241 次
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 被引用 213 次
- Share or Not? Learning to Schedule Language-Specific Capacity for Multilingual TranslationBiao Zhang, Ankur Bapna, Rico Sennrich, Orhan FiratICLR 2021 · 被引用 97 次
- Parameter-Efficient Fine-Tuning without Introducing New LatencyBaohao Liao, Yan Meng, Christof MonzACL 2023 · 被引用 26 次
相关 Paper
- Parameter Differentiation Based Multilingual Neural Machine TranslationQian Wang, Jiajun ZhangAAAI 2022 · 被引用 21 次
- Importance-based Neuron Allocation for Multilingual Neural Machine TranslationWanying Xie, Yang Feng, Shuhao Gu, Dong YuACL 2021
- Interpreting and Exploiting Functional Specialization in Multi-Head Attention under Multi-task LearningChong Li, Shaonan Wang, Yunhao Zhang, Jiajun Zhang 等EMNLP 2023 · 被引用 5 次
- On Negative Interference in Multilingual Models: Findings and A Meta-Learning TreatmentZirui Wang, Zachary C. Lipton, Yulia TsvetkovEMNLP 2020 · 被引用 72 次
- LANDeRMT: Dectecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine TranslationShaolin Zhu, Leiyu Pan, Bo Li, Deyi XiongACL 2024
