Understanding MLP-Mixer as a wide and sparse MLP
Tomohiro Hayase, Ryo Karakida
Abstract
Multi-layer perceptron (MLP) is a fundamental component of deep learning, and recent MLP-based architectures, especially the MLP-Mixer, have achieved significant empirical success. Nevertheless, our understanding of why and how the MLP-Mixer outperforms conventional MLPs remains largely unexplored. In this work, we reveal that sparseness is a key mechanism underlying the MLP-Mixers. First, the Mixers have an effective expression as a wider MLP with Kronecker-product weights, clarifying that the Mixers efficiently embody several sparseness properties explored in deep learning. In the case of linear layers, the effective expression elucidates an implicit sparse regularization caused by the model architecture and a hidden relation to Monarch matrices, which is also known as another form of sparse parameterization. Next, for general cases, we empirically demonstrate quantitative similarities between the Mixer and the unstructured sparse-weight MLPs. Following a guiding principle proposed by Golubeva, Neyshabur and Gur-Ari (2021), which fixes the number of connections and increases the width and sparsity, the Mixers can demonstrate improved performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af2e4f2f-aec7-4bd5-bb81-b1abb8972230Cited by top-tier papers2
- Beyond Dense Connectivity: Explicit Sparsity for Scalable RecommendationYantao Yu, Sen Qiao, Lei Shen, Bing Wang et al.SIGIR 2026 · 3 citations
- Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural NetworksJieyuan Zhang, Xiaolong Zhou, Shuai Wang, Wenjie Wei et al.NeurIPS 2025 · 2 citations
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si et al.CVPR 2022 · 1,114 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
Related papers
- Searching for Efficient Linear Layers over a Continuous Space of Structured MatricesAndres Potapczynski, Shikai Qiu, Marc Finzi, Christopher Ferri et al.NeurIPS 2024 · 11 citations
- Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of DecodersJames Oldfield, Shawn Im, Sharon Li, Mihalis A. Nicolaou et al.NeurIPS 2025 · 7 citations
- DynaMixer: A Vision MLP Architecture with Dynamic MixingZiyu Wang, Wenhao Jiang, Yiming Zhu, Li Yuan et al.ICML 2022 · 55 citations
- QiMLP: Quantum-inspired Multilayer Perceptron with Strong Correlation Mining and Parameter CompressionJunwei Zhang, Tianheng Wang, Zeyi Zhang, Pengju Yan et al.AAAI 2025 · 1 citation
- Towards Understanding the Mixture-of-Experts Layer in Deep LearningZixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu et al.NeurIPS 2022 · 199 citations
