Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts
Basil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton, Neil Houlsby
Abstract
Large sparsely-activated models have obtained excellent performance in multiple domains. However, such models are typically trained on a single modality at a time. We present the Language-Image MoE, LIMoE, a sparse mixture of experts model capable of multimodal learning. LIMoE accepts both images and text simultaneously, while being trained using a contrastive loss. MoEs are a natural fit for a multimodal backbone, since expert layers can learn an appropriate partitioning of modalities. However, new challenges arise; in particular, training stability and balanced expert utilization, for which we propose an entropy-based regularization scheme. Across multiple scales, we demonstrate remarkable performance improvement over dense models of equivalent computational cost. LIMoE-L/16 trained comparably to CLIP-L/14 achieves 78.6% zero-shot ImageNet accuracy (vs. 76.2%), and when further scaled to H/14 (with additional data) it achieves 84.1%, comparable to state-of-the-art methods which use larger custom per-modality backbones and pre-training schemes. We analyse the quantitative and qualitative behavior of LIMoE, and demonstrate phenomena such as differing treatment of the modalities and the organic emergence of modality-specific experts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01499e68-5d08-4bab-b977-1833379b2abfCited by top-tier papers106
- From Sparse to Soft Mixtures of ExpertsJoan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, Neil HoulsbyICLR 2024 · 264 citations
- OpenMoE: An Early Effort on Open Mixture-of-Experts Language ModelsFuzhao Xue, Zian Zheng, Yao Fu, Jinjie Ni et al.ICML 2024 · 183 citations
- Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction TuningTed Zadouri, Ahmet Üstün, Arash Ahmadian, Beyza Ermis et al.ICLR 2024 · 169 citations
- ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich ManipulationJiawen Yu, Hairuo Liu, Qiaojun Yu, Jieji Ren et al.NeurIPS 2025 · 150 citations
- FuseMoE: Mixture-of-Experts Transformers for Fleximodal FusionXing Han, Huy Nguyen, Carl Harris, Nhat Ho et al.NeurIPS 2024 · 129 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
Related papers
- CLIP-FMoE: Scalable CLIP via Fused Mixture-of-Experts with Enforced SpecializationLuong Tran, Lan-Cuong Nguyen, Huynh Dang Nguyen, Dat Nguyen-Cong et al.ICLR 2026
- MoDE: CLIP Data Experts via ClusteringJiawei Ma, Po-Yao Huang, Saining Xie, Shang-Wen Li et al.CVPR 2024 · 8 citations
- Soft Modality-Guided Expert Specialization in MoE-VLMsZi-Hao Bo, Yaqian Li, Anzhou Hou, Rinyoichi Takezoe et al.CVPR 2026
- CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet UpcyclingJihai Zhang, Xiaoye Qu, Tong Zhu, Yu ChengEMNLP 2025 · 1 citation
- ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable SpecializationAnzhe Cheng, Shukai Duan, Shixuan Li, Chenzhong Yin et al.CVPR 2026 · 8 citations
