UMoE: Unifying Attention and FFN with Shared Experts
Yuanhang Yang, Chaozheng Wang, Jing Li
Abstract
Sparse Mixture of Experts (MoE) architectures have emerged as a promising approach for scaling Transformer models. While initial works primarily incorporated MoE into feed-forward network (FFN) layers, recent studies have explored extending the MoE paradigm to attention layers to enhance model performance. However, existing attention-based MoE layers require specialized implementations and demonstrate suboptimal performance compared to their FFN-based counterparts. In this paper, we aim to unify MoE designs in attention and FFN layers by introducing a novel reformulation of the attention mechanism, that reveals an underlying FFN-like structure within attention modules. Our proposed architecture, UMoE, achieves superior performance through attention-based MoE layers while enabling efficient parameter sharing between FFN and attention components.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94076e79-b1b0-4d82-919a-ceb603ac920fCited by top-tier papers1
Ask how each one uses itBuilds on20
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang et al.ICML 2022 · 523 citations
- Parallelizing Linear Transformers with the Delta Rule over Sequence LengthSonglin Yang, Bailin Wang, Yu Zhang, Yikang Shen et al.NeurIPS 2024 · 412 citations
- BASE Layers: Simplifying Training of Large, Sparse ModelsMike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal et al.ICML 2021 · 382 citations
Related papers
- MoEUT: Mixture-of-Experts Universal TransformersRóbert Csordás, Kazuki Irie, Jürgen Schmidhuber, Christopher Potts et al.NeurIPS 2024 · 65 citations
- BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of ExpertsQizhen (Irene) Zhang, Nikolas Gritsch, Dwaraknath Gnaneshwar, Simon Guo et al.NeurIPS 2024 · 18 citations
- UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context LearningZihao Huang, Yu Bao, Qiyang Min, Siyan Chen et al.ICLR 2026 · 6 citations
- Mixture of Attention Heads: Selecting Attention Heads Per TokenXiaofeng Zhang, Yikang Shen, Zeyu Huang, Jie Zhou et al.EMNLP 2022 · 23 citations
- Ultra-Sparse Memory NetworkZihao Huang, Qiyang Min, Hongzhi Huang, Yutao Zeng et al.ICLR 2025
