Autonomy-of-Experts Models
Ang Lv, Ruobing Xie, Yining Qian, Songhao Wu, Xingwu Sun, Zhanhui Kang, Di Wang, Rui Yan
Abstract
Mixture-of-Experts (MoE) models mostly use a router to assign tokens to specific expert modules, activating only partial parameters and often outperforming dense models. We argue that the separation between the router's decision-making and the experts' execution is a critical yet overlooked issue, leading to suboptimal expert selection and ineffective learning. To address this, we propose Autonomy-of-Experts (AoE), a novel MoE paradigm in which experts autonomously select themselves to process inputs. AoE is based on the insight that an expert is aware of its own capacity to effectively process a token, an awareness reflected in the scale of its internal activations. In AoE, routers are removed; instead, experts pre-compute internal activations for inputs and are ranked based on their activation norms. Only the top-ranking experts proceed with the forward pass, while the others abort. The overhead of pre-computing activations is reduced through a low-rank weight factorization. This self-evaluating-then-partner-comparing approach ensures improved expert selection and effective learning. We pre-train language models having 700M up to 4B parameters, demonstrating that AoE outperforms traditional MoE models with comparable efficiency. The code is available at https://github.com/trestad/ Autonomy-of-Experts
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24470f88-8ce4-4cd6-bccf-dd59b187e4ddCited by top-tier papers3
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-ExpertsHeming Zou, Yunliang Zang, Wutong Xu, Yao Zhu et al.NeurIPS 2025 · 38 citations
- Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary LossAng Lv, Jin Ma, Yiyuan Ma, Siyuan QiaoICLR 2026 · 14 citations
- TIME: Tensor-Factorized Mixture-of-Experts with Intrinsic Routing for Lifelong Multimodal Knowledge EditingDexuan Xu, Jieyi Wang, Shijie Li, Hanpin Wang et al.ICML 2026
Builds on15
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- BASE Layers: Simplifying Training of Large, Sparse ModelsMike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal et al.ICML 2021 · 382 citations
- Hash Layers For Large Sparse ModelsStephen Roller, Sainbayar Sukhbaatar, Arthur Szlam, Jason WestonNeurIPS 2021 · 316 citations
Related papers
- Union-of-Experts: Neurons in Mixture-of-Experts are Secretly RoutersSonghao Wu, Ang Lv, Ruobing Xie, Samm Sun et al.ACL 2026
- HMoE: Heterogeneous Mixture of Experts for Language ModelingAn Wang, Xingwu Sun, Ruobing Xie, Shuaipeng Li et al.EMNLP 2025 · 2 citations
- Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer ModelsYongxin Guo, Zhenglin Cheng, Xiaoying Tang, Zhaopeng Tu et al.ICLR 2025
- Harder Task Needs More Experts: Dynamic Routing in MoE ModelsQuzhe Huang, Zhenwei An, Nan Zhuang, Mingxu Tao et al.ACL 2024 · 11 citations
- Layerwise Recurrent Router for Mixture-of-ExpertsZihan Qiu, Zeyu Huang, Shuang Cheng, Yizhi Zhou et al.ICLR 2025
