TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice
Shen Yan, Xingyan Bin, Sijun Zhang, Yisen Wang, Zhouchen Lin
Abstract
The Mixture of Experts (MoE) architecture has emerged as a promising solution to reduce computational overhead by selectively activating subsets of model parameters. The effectiveness of MoE models depends primarily on their routing mechanisms, with the widely adopted Top-K routing scheme used for activating experts. However, the Top-K scheme has notable limitations, including unnecessary activations and underutilization of experts. In this work, rather than modifying the routing mechanism as done in previous studies, we propose the Ternary Choice MoE (TC-MoE), a novel approach that expands the expert space by applying the ternary set -1, 0, 1 to each expert. This expansion allows more efficient and effective expert activations without incurring significant computational costs. Additionally, given the unique characteristics of the expanded expert space, we introduce a new load balance loss and reward loss to ensure workload balance and achieve a flexible trade-off between effectiveness and efficiency. Extensive experiments demonstrate that TC-MoE achieves an average improvement of over 1.1% compared with traditional approaches, while reducing the average number of activated experts by up to 9%. These results confirm that TC-MoE effectively addresses the inefficiencies of conventional routing schemes, offering a more efficient and scalable solution for MoE-based large language models. Code and models are available at https://github.com/stiger1000/TC-MoE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c96957d1-5ec7-4a7d-ab4b-d0c59f9666c4Cited by top-tier papers3
- Grouter: Decoupling Routing from Representation for Accelerated MoE TrainingYuqi Xu, Rizhen Hu, zihan liu, Mou Sun et al.ICML 2026 · 7 citations
- Certain Head, Uncertain Tail: Expert-Sample for Test-Time Scaling in Fine-Grained MoEYuanteng Chen, Peisong Wang, Nanxin Zeng, Yuantian Shao et al.ICML 2026 · 3 citations
- Mixture of Lookup ExpertsShibo Jie, Yehui Tang, Kai Han, Yitong Li et al.ICML 2025
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
Related papers
- Harder Task Needs More Experts: Dynamic Routing in MoE ModelsQuzhe Huang, Zhenwei An, Nan Zhuang, Mingxu Tao et al.ACL 2024 · 11 citations
- Autonomy-of-Experts ModelsAng Lv, Ruobing Xie, Yining Qian, Songhao Wu et al.ICML 2025
- Ada-K Routing: Boosting the Efficiency of MoE-based LLMsTongtian Yue, Longteng Guo, Jie Cheng, Xuange Gao et al.ICLR 2025
- Advancing Expert Specialization for Better MoEHongcan Guo, Haolang Lu, Guoshun Nan, Bolun Chu et al.NeurIPS 2025 · 48 citations
- HMoE: Heterogeneous Mixture of Experts for Language ModelingAn Wang, Xingwu Sun, Ruobing Xie, Shuaipeng Li et al.EMNLP 2025 · 2 citations
