Ultra-Sparse Memory Network
Zihao Huang, Qiyang Min, Hongzhi Huang, Yutao Zeng, Defa Zhu, Ran Guo, Xun Zhou
Abstract
It is widely acknowledged that the performance of Transformer models is logarithmically related to their number of parameters and computational complexity. While approaches like Mixture of Experts (MoE) decouple parameter count from computational complexity, they still face challenges in inference due to high memory access costs. This work introduces UltraMem, incorporating large-scale, ultrasparse memory layer to address these limitations. Our approach significantly reduces inference latency while maintaining model performance. We also investigate the scaling laws of this new architecture, demonstrating that it not only exhibits favorable scaling properties but outperforms MoE. In experiments, the largest UltraMem we train has 20 million memory slots. The results show that our method achieves state-of-the-art inference speed and model performance within a given computational budget, paving the way for billions of slots or experts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2fff2cdb-f448-4fe2-8e42-adb27e41413aCited by top-tier papers10
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language ModelsXin Cheng, Wangding Zeng, Damai Dai, Qinyu Chen et al.ACL 2026 · 57 citations
- SonicMoE: Accelerating MoE with IO and Tile-aware OptimizationsWentao Guo, Mayank Mishra, Xinle Cheng, Ion Stoica et al.ICLR 2026 · 22 citations
- STEM: Scaling Transformers with Embedding ModulesRanajoy Sadhukhan, Sheng Cao, Harry Dong, Changsheng Zhao et al.ICLR 2026 · 14 citations
- Pretraining with hierarchical memories: separating long-tail and common knowledgeHadi Pouransari, David Grangier, C Thomas, Michael Kirchhof et al.ICLR 2026 · 11 citations
- UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context LearningZihao Huang, Yu Bao, Qiyang Min, Siyan Chen et al.ICLR 2026 · 6 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
Related papers
- UMoE: Unifying Attention and FFN with Shared ExpertsYuanhang Yang, Chaozheng Wang, Jing LiNeurIPS 2025 · 4 citations
- Exploiting Activation Sparsity with Dense to Dynamic-k Mixture-of-Experts ConversionFilip Szatkowski, Bartosz Wójcik, Mikolaj Piórczynski, Simone ScardapaneNeurIPS 2024 · 19 citations
- Sparse Universal TransformerShawn Tan, Yikang Shen, Zhenfang Chen, Aaron C. Courville et al.EMNLP 2023 · 6 citations
- Hyperparameter Transfer with Mixture-of-Expert LayersTianze Jiang, Blake Bordelon, Cengiz Pehlevan, Boris HaninICML 2026 · 6 citations
- Mixture of Parrots: Experts improve memorization more than reasoningSamy Jelassi, Clara Mohri, David Brandfonbrener, Alex Gu et al.ICLR 2025
