Mixture of Hidden-Dimensions: Not All Hidden-States' Dimensions are Needed in Transformer
Yilong Chen, Junyuan Shang, Zhenyu Zhang, Jiawei Sheng, Tingwen Liu, Shuohuan Wang, Yu Sun, Hua Wu, Haifeng Wang
Abstract
Transformer models encounter inefficiency when scaling hidden dimensions due to the uniform expansion of parameters. When delving into the sparsity of hidden dimensions, we observe that only a small subset of dimensions are highly activated, where some dimensions are commonly activated across tokens, and some others uniquely activated for individual tokens. To leverage this, we propose MOHD (Mixture of Hidden Dimensions), a sparse architecture that combines shared sub-dimensions for common features and dynamically routes specialized sub-dimensions per token. To address the potential information loss from sparsity, we introduce activation scaling and group fusion mechanisms. MOHD efficiently expands hidden dimensions with minimal computational increases, outperforming vanilla Transformers in both parameter efficiency and task performance across 10 NLP tasks. MOHD achieves 1.7% higher performance with 50% fewer activated parameters and 3.7% higher performance with 3× total parameters expansion at constant activated parameters cost. MOHD offers a new perspective for scaling the model, showcasing the potential of hidden dimension sparsity. … MoHD Multi-Head Attention RMSNorm MoHD Feed Forward RMSNorm Layer L MoHD Blocks
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7f731a5-b922-40fb-a9e8-506ae2679b33Builds on20
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- Deja Vu: Contextual Sparsity for Efficient LLMs at Inference TimeZichang Liu, Jue Wang, Tri Dao, Tianyi Zhou et al.ICML 2023 · 318 citations
Related papers
- Exploiting Activation Sparsity with Dense to Dynamic-k Mixture-of-Experts ConversionFilip Szatkowski, Bartosz Wójcik, Mikolaj Piórczynski, Simone ScardapaneNeurIPS 2024 · 19 citations
- MoH: Multi-Head Attention as Mixture-of-Head AttentionPeng Jin, Bo Zhu, Li Yuan, Shuicheng YanICML 2025 · 2 citations
- Hyperparameter Transfer with Mixture-of-Expert LayersTianze Jiang, Blake Bordelon, Cengiz Pehlevan, Boris HaninICML 2026 · 6 citations
- Exploring the Benefit of Activation Sparsity in Pre-trainingZhengyan Zhang, Chaojun Xiao, Qiujieli Qin, Yankai Lin et al.ICML 2024 · 6 citations
- Mixture of Attention Heads: Selecting Attention Heads Per TokenXiaofeng Zhang, Yikang Shen, Zeyu Huang, Jie Zhou et al.EMNLP 2022 · 23 citations
