A Theoretical View on Sparsely Activated Networks
Cenk Baykal, Nishanth Dikkala, Rina Panigrahy, Cyrus Rashtchian, Xin Wang
Abstract
Deep and wide neural networks successfully fit very complex functions today, but dense models are starting to be prohibitively expensive for inference. To mitigate this, one promising direction is networks that activate a sparse subgraph of the network. The subgraph is chosen by a data-dependent routing function, enforcing a fixed mapping of inputs to subnetworks (e.g., the Mixture of Experts (MoE) paradigm in Switch Transformers). However, prior work is largely empirical, and while existing routing functions work well in practice, they do not lead to theoretical guarantees on approximation ability. We aim to provide a theoretical explanation for the power of sparse networks. As our first contribution, we present a formal model of data-dependent sparse networks that captures salient aspects of popular architectures. We then introduce a routing function based on locality sensitive hashing (LSH) that enables us to reason about how well sparse networks approximate target functions. After representing LSH-based sparse networks with our model, we prove that sparse networks can match the approximation power of dense networks on Lipschitz functions. Applying LSH on the input vectors means that the experts interpolate the target function in different subregions of the input space. To support our theory, we define various datasets based on Lipschitz target functions, and we show that sparse networks give a favorable trade-off between number of active units and approximation quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88a9a8a6-a9ac-4c5c-b1be-0e40799cc35aCited by top-tier papers4
- Alternating Updates for Efficient TransformersCenk Baykal, Dylan J. Cutler, Nishanth Dikkala, Nikhil Ghosh et al.NeurIPS 2023 · 13 citations
- The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in TransformersZonglin Li, Chong You, Srinadh Bhojanapalli, Daliang Li et al.ICLR 2023 · 10 citations
- On the Benefits of Learning to Route in Mixture-of-Experts ModelsNishanth Dikkala, Nikhil Ghosh, Raghu Meka, Rina Panigrahy et al.EMNLP 2023 · 9 citations
- On the Expressive Power of Mixture-of-Experts for Structured Complex TasksMingze Wang, Weinan ENeurIPS 2025 · 3 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Hash Layers For Large Sparse ModelsStephen Roller, Sainbayar Sukhbaatar, Arthur Szlam, Jason WestonNeurIPS 2021 · 316 citations
Related papers
- On the Adversarial Robustness of Mixture of ExpertsJoan Puigcerver, Rodolphe Jenatton, Carlos Riquelme, Pranjal Awasthi et al.NeurIPS 2022 · 33 citations
- Towards Understanding the Mixture-of-Experts Layer in Deep LearningZixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu et al.NeurIPS 2022 · 199 citations
- LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive HashingXiaonan Nie, Qibin Liu, Fangcheng Fu, Shenhan Zhu et al.NeurIPS 2024 · 10 citations
- RouterInterp: Understanding Superposed Specialisation in Mixture of Experts RoutingIlya Lasy, Nora Cai, Kola AyonrindeICML 2026
- Efficient Quantization of Mixture-of-Experts with Theoretical Generalization GuaranteesMohammed Nowaz Rabbani Chowdhury, Kaoutar El Maghraoui, Hsinyu Tsai, Naigang Wang et al.ICLR 2026 · 2 citations
