MoDE: CLIP Data Experts via Clustering
Jiawei Ma, Po-Yao Huang, Saining Xie, Shang-Wen Li, Luke Zettlemoyer, Shih-Fu Chang, Wen-Tau Yih, Hu Xu
Abstract
The success of contrastive language-image pretraining (CLIP) relies on the supervision from the pairing between images and captions, which tends to be noisy in webcrawled data. We present Mixture of Data Experts (MoDE) and learn a system of CLIP data experts via clustering. Each data expert is trained on one data cluster, being less sensitive to false negative noises in other clusters. At inference time, we ensemble their outputs by applying weights determined through the correlation between task metadata and cluster conditions. To estimate the correlation precisely, the samples in one cluster should be semantically similar, but the number of data experts should still be reasonable for training and inference. As such, we consider the ontology in human language and propose to use finegrained cluster centers to represent each data expert at a coarse-grained level. Experimental studies show that four CLIP data experts on ViT-B/16 outperform the ViT-L/14 by OpenAI CLIP and OpenCLIP on zero-shot image classification but with less (<35%) training cost. Meanwhile, MoDE can train all data expert asynchronously and can flexibly include new data experts. The code is available here. * Research done while Jiawei Ma was an intern at FAIR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7ff7623-a24c-43f9-b35b-d33f1ead232dCited by top-tier papers13
- CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic SegmentationDengke Zhang, Fagui Liu, Quan TangICCV 2025 · 6 citations
- LogiCoL: Logically-Informed Contrastive Learning for Set-based Dense RetrievalYanzhen Shen, Sihao Chen, Xueqiang Xu, Yunyi Zhang et al.EMNLP 2025 · 1 citation
- CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet UpcyclingJihai Zhang, Xiaoye Qu, Tong Zhu, Yu ChengEMNLP 2025 · 1 citation
- Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image ClassificationZequn Zeng, Yudi Su, Jianqiao Sun, Tiansheng Wen et al.CVPR 2025
- Mixing Expertise with Confidence: A Mixture of Experts Framework for Robust Multi-Modal Continual LearningMd Abdullah Al Forhad, Yuansheng Zhu Zhu., Abhinab Acharya, Xumin Liu et al.ICML 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
Related papers
- CLIP-FMoE: Scalable CLIP via Fused Mixture-of-Experts with Enforced SpecializationLuong Tran, Lan-Cuong Nguyen, Huynh Dang Nguyen, Dat Nguyen-Cong et al.ICLR 2026
- Demystifying CLIP DataHu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang et al.ICLR 2024 · 249 citations
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training ParadigmYangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui et al.ICLR 2022 · 565 citations
- Non-Contrastive Learning Meets Language-Image Pre-TrainingJinghao Zhou, Li Dong, Zhe Gan, Lijuan Wang et al.CVPR 2023
- Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual KnowledgeFawaz Sammani, Nikos DeligiannisNeurIPS 2024 · 15 citations
