Mixture-of-Modules: Reinventing Transformers as Dynamic Assemblies of Modules
Zhuocheng Gong, Ang Lv, Jian Guan, Wei Wu, Huishuai Zhang, Minlie Huang, Dongyan Zhao, Rui Yan
Abstract
Is it always necessary to compute tokens from shallow to deep layers in Transformers? The continued success of vanilla Transformers and their variants suggests an undoubted "yes". In this work, however, we attempt to break the depthordered convention by proposing a novel architecture dubbed mixture-of-modules (MoM), which is motivated by an intuition that any layer, regardless of its position, can be used to compute a token as long as it possesses the needed processing capabilities. The construction of MoM starts from a finite set of modules defined by multi-head attention and feed-forward networks, each distinguished by its unique parameterization. Two routers then iteratively select attention modules and feed-forward modules from the set to process a token. The selection dynamically expands the computation graph in the forward pass of the token, culminating in an assembly of modules. We show that MoM provides not only a unified framework for Transformers and their numerous variants but also a flexible and learnable approach for reducing redundancy in Transformer parameterization. We pre-train various MoMs using OpenWebText. Empirical results demonstrate that MoMs, of different parameter counts, consistently outperform vanilla transformers on both GLUE and XSUM benchmarks. More interestingly, with a fixed parameter budget, MoM-large enables an over 38% increase in depth for computation graphs compared to GPT-2-large, resulting in absolute gains of 1.4 on GLUE and 1 on XSUM. On the other hand, MoM-large also enables an over 60% reduction in depth while involving more modules per layer, yielding a 16% reduction in TFLOPs and a 43% decrease in memory usage compared to GPT-2-large, while maintaining comparable performance. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b8533a3-9dea-46ba-b981-e50f5f1a5744Cited by top-tier papers7
- Mixture of In-Context Experts Enhance LLMs' Long Context AwarenessHongzhan Lin, Ang Lv, Yuhan Chen, Chen Zhu et al.NeurIPS 2024 · 25 citations
- Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary LossAng Lv, Jin Ma, Yiyuan Ma, Siyuan QiaoICLR 2026 · 14 citations
- D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM ServingHaodong Wang, Qihua Zhou, Zicong Hong, Song GuoMobiCom 2025 · 8 citations
- PEAR: Position-Embedding-Agnostic Attention Re-weighting Enhances Retrieval-Augmented Generation with Zero Inference OverheadTao Tan, Yining Qian, Ang Lv, Hongzhan Lin et al.WWW 2025 · 4 citations
- Autonomy-of-Experts ModelsAng Lv, Ruobing Xie, Yining Qian, Songhao Wu et al.ICML 2025
Builds on10
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley et al.NeurIPS 2020 · 473 citations
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
- Depth-Adaptive TransformerMaha Elbayad, Jiatao Gu, Edouard Grave, Michael AuliICLR 2020 · 264 citations
Related papers
- Mixture of Tokens: Continuous MoE through Cross-Example AggregationSzymon Antoniak, Michal Krutul, Maciej Pióro, Jakub Krajewski et al.NeurIPS 2024 · 6 citations
- Hyperparameter Transfer with Mixture-of-Expert LayersTianze Jiang, Blake Bordelon, Cengiz Pehlevan, Boris HaninICML 2026 · 6 citations
- MoEUT: Mixture-of-Experts Universal TransformersRóbert Csordás, Kazuki Irie, Jürgen Schmidhuber, Christopher Potts et al.NeurIPS 2024 · 65 citations
- MoM: Linear Sequence Modeling with Mixture-of-MemoriesJusen Du, Weigao Sun, Disen Lan, Jiaxi Hu et al.ICLR 2026 · 43 citations
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech RecognitionUmberto Cappellazzo, Minsu Kim, Pingchuan Ma, Honglie Chen et al.NeurIPS 2025 · 5 citations
