MoEUT: Mixture-of-Experts Universal Transformers
Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber, Christopher Potts, Christopher D. Manning
Abstract
Previous work on Universal Transformers (UTs) has demonstrated the importance of parameter sharing across layers. By allowing recurrence in depth, UTs have advantages over standard Transformers in learning compositional generalizations, but layer-sharing comes with a practical limitation of parameter-compute ratio: it drastically reduces the parameter count compared to the non-shared model with the same dimensionality. Naively scaling up the layer size to compensate for the loss of parameters makes its computational resource requirements prohibitive. In practice, no previous work has succeeded in proposing a shared-layer Transformer design that is competitive in parameter count-dominated tasks such as language modeling. Here we propose MoEUT (pronounced"moot"), an effective mixture-of-experts (MoE)-based shared-layer Transformer architecture, which combines several recent advances in MoEs for both feedforward and attention layers of standard Transformers together with novel layer-normalization and grouping schemes that are specific and crucial to UTs. The resulting UT model, for the first time, slightly outperforms standard Transformers on language modeling tasks such as BLiMP and PIQA, while using significantly less compute and memory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 56ff7722-8604-4548-b9ac-d83987cda321Cited by top-tier papers16
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachJonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer et al.NeurIPS 2025 · 431 citations
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeZihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang et al.NeurIPS 2025 · 336 citations
- Do Language Models Use Their Depth Efficiently?Róbert Csordás, Christopher D. Manning, Christopher PottsNeurIPS 2025 · 61 citations
- Improved Representation Steering for Language ModelsZhengxuan Wu, Qinan Yu, Aryaman Arora, Christopher D. Manning et al.NeurIPS 2025 · 22 citations
- MeSH: Memory-as-State-Highways for Recursive TransformersChengting Yu, Xiaobo Shu, Yadao Wang, Yizhen Zhang et al.ICLR 2026 · 10 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- Sparse Universal TransformerShawn Tan, Yikang Shen, Zhenfang Chen, Aaron C. Courville et al.EMNLP 2023 · 6 citations
- UMoE: Unifying Attention and FFN with Shared ExpertsYuanhang Yang, Chaozheng Wang, Jing LiNeurIPS 2025 · 4 citations
- SwitchHead: Accelerating Transformers with Mixture-of-Experts AttentionRóbert Csordás, Piotr Piekos, Kazuki Irie, Jürgen SchmidhuberNeurIPS 2024 · 46 citations
- BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of ExpertsQizhen (Irene) Zhang, Nikolas Gritsch, Dwaraknath Gnaneshwar, Simon Guo et al.NeurIPS 2024 · 18 citations
- Layerwise Recurrent Router for Mixture-of-ExpertsZihan Qiu, Zeyu Huang, Shuang Cheng, Yizhi Zhou et al.ICLR 2025
