THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine Translation
Yunlong Liang, Fandong Meng, Jie Zhou
Abstract
The sparse Mixture-of-Experts (MoE) has achieved significant progress for neural machine translation (NMT). However, there exist two limitations in current MoE solutions which may lead to sub-optimal performance: 1) they directly use the task knowledge of NMT into MoE (e.g., domain/linguistics-specific knowledge), which are generally unavailable at practical application and neglect the naturally grouped domain/linguistic properties; 2) the expert selection only depends on the localized token representation without considering the context, which fully grasps the state of each token in a global view. To address the above limitations, we propose THOR-MoE via arming the MoE with hierarchical task-guided and context-responsive routing policies. Specifically, it 1) firstly predicts the domain/language label and then extracts mixed domain/language representation to allocate task-level experts in a hierarchical manner; 2) injects the context information to enhance the token routing from the pre-selected task-level experts set, which can help each token to be accurately routed to more specialized and suitable experts. Extensive experiments on multi-domain translation and multilingual translation benchmarks with different architectures consistently demonstrate the superior performance of THOR-MoE. Additionally, the THOR-MoE operates as a plug-and-play module compatible with existing Top- and Top- routing schemes, ensuring broad applicability across diverse MoE architectures. For instance, compared with vanilla Top- routing, the context-aware manner can achieve an average improvement of 0.75 BLEU with less than 22% activated parameters on multi-domain translation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f4b9ff3-9ebd-4ff1-9433-4eecffe00dcdCited by top-tier papers1
Ask how each one uses itBuilds on16
- Better & Faster Large Language Models via Multi-token PredictionFabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz et al.ICML 2024 · 286 citations
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 213 citations
- Taming Sparsely Activated Transformer with Stochastic ExpertsSimiao Zuo, Xiaodong Liu, Jian Jiao, Young Jin Kim et al.ICLR 2022 · 144 citations
- Share or Not? Learning to Schedule Language-Specific Capacity for Multilingual TranslationBiao Zhang, Ankur Bapna, Rico Sennrich, Orhan FiratICLR 2021 · 97 citations
- Improving Multilingual Translation by Representation and Gradient RegularizationYilin Yang, Akiko Eriguchi, Alexandre Muzio, Prasad Tadepalli et al.EMNLP 2021 · 16 citations
Related papers
- HyperMoE: Towards Better Mixture of Experts via Transferring Among ExpertsHao Zhao, Zihan Qiu, Huijia Wu, Zili Wang et al.ACL 2024
- MMNMT: Modularizing Multilingual Neural Machine Translation with Flexibly Assembled MoE and Dense BlocksShangjie Li, Xiangpeng Wei, Shaolin Zhu, Jun Xie et al.EMNLP 2023 · 4 citations
- Multi-Head Mixture-of-ExpertsXun Wu, Shaohan Huang, Wenhui Wang, Shuming Ma et al.NeurIPS 2024 · 42 citations
- How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-ExpertsSumin Park, Noseong ParkAAAI 2026
- Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of ExpertsYongXiang Hua, Haoyu Cao, Zhou Tao, Bocheng Li et al.ACM MM 2025 · 1 citation
