DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning
Hussein Hazimeh, Zhe Zhao, Aakanksha Chowdhery, Maheswaran Sathiamoorthy, Yihua Chen, Rahul Mazumder, Lichan Hong, Ed H. Chi
摘要
The Mixture-of-Experts (MoE) architecture is showing promising results in improving parameter sharing in multi-task learning (MTL) and in scaling high-capacity neural networks. State-of-the-art MoE models use a trainable sparse gate to select a subset of the experts for each input example. While conceptually appealing, existing sparse gates, such as Top-k, are not smooth. The lack of smoothness can lead to convergence and statistical performance issues when training with gradient-based methods. In this paper, we develop DSelect-k: a continuously differentiable and sparse gate for MoE, based on a novel binary encoding formulation. The gate can be trained using first-order methods, such as stochastic gradient descent, and offers explicit control over the number of experts to select. We demonstrate the effectiveness of DSelect-k on both synthetic and real MTL datasets with up to tasks. Our experiments indicate that DSelect-k can achieve statistically significant improvements in prediction and expert selection over popular MoE gates. Notably, on a real-world, large-scale recommender system, DSelect-k achieves over improvement in predictive performance compared to Top-k. We provide an open-source implementation of DSelect-k.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper56
- Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of ExpertsBasil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton 等NeurIPS 2022 · 被引用 359 次
- Unified Scaling Laws for Routed Language ModelsAidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch 等ICML 2022 · 被引用 266 次
- AdaMerging: Adaptive Model Merging for Multi-Task LearningEnneng Yang, Zhenyi Wang, Li Shen, Shiwei Liu 等ICLR 2024 · 被引用 230 次
- Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction TuningTed Zadouri, Ahmet Üstün, Arash Ahmadian, Beyza Ermis 等ICLR 2024 · 被引用 169 次
- M³ViT: Mixture-of-Experts Vision Transformer for Efficient Multi-task Learning with Model-Accelerator Co-designHanxue Liang, Zhiwen Fan, Rishov Sarkar, Ziyu Jiang 等NeurIPS 2022 · 被引用 152 次
它引用的顶会 Paper5
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- AdaShare: Learning What To Share For Efficient Deep Multi-Task LearningXimeng Sun, Rameswar Panda, Rogério Feris, Kate SaenkoNeurIPS 2020 · 被引用 337 次
- Differentiable Top-k with Optimal TransportYujia Xie, Hanjun Dai, Minshuo Chen, Bo Dai 等NeurIPS 2020 · 被引用 124 次
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause 等NeurIPS 2020 · 被引用 104 次
- The Tree Ensemble Layer: Differentiability meets Conditional ComputationHussein Hazimeh, Natalia Ponomareva, Petros Mol, Zhenyu Tan 等ICML 2020 · 被引用 95 次
相关 Paper
- COMET: Learning Cardinality Constrained Mixture of Experts with Trees and Local SearchShibal Ibrahim, Wenyu Chen, Hussein Hazimeh, Natalia Ponomareva 等KDD 2023 · 被引用 2 次
- ReMoE: Fully Differentiable Mixture-of-Experts with ReLU RoutingZiteng Wang, Jun Zhu, Jianfei ChenICLR 2025
- Theory on Mixture-of-Experts in Continual LearningHongbo Li, Sen Lin, Lingjie Duan, Yingbin Liang 等ICLR 2025
- ProbMoE: Differentiable Probabilistic Routing for Mixture-of-ExpertsHeng Zhao, Zilei Shao, Guy Van den Broeck, Zhe ZengICML 2026
- SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMsMikołaj Zasada, Łukasz Struski, Jacek Tabor, Marcin KurdzielICML 2026 · 被引用 2 次
