CLIP-FMoE: Scalable CLIP via Fused Mixture-of-Experts with Enforced Specialization
Luong Tran, Lan-Cuong Nguyen, Huynh Dang Nguyen, Dat Nguyen-Cong, Dung D. Le, Van Nguyen
摘要
Mixture-of-Experts (MoE) architectures have emerged as a promising approach for scaling deep learning models while maintaining computational efficiency. However, existing MoE adaptations for Contrastive Language-Image Pre-training (CLIP) models suffer from significant computational overhead during sequential training and degradation of zero-shot capabilities. To address these limitations, we propose CLIP-FMoE, a novel approach that integrates MoE architecture into CLIP fine-tuning. Our method uses Isolated Constrained Contrastive Learning, a pipeline that trains specialized experts on cluster-based data partitions to accelerate expert specialization. Additionally, we introduce a Fusion Gate mechanism to mitigate catastrophic forgetting of pre-trained knowledge. Extensive experiments across multiple benchmarks demonstrate that our approach achieves consistent improvements on downstream tasks while preserving zero-shot capabilities. Furthermore, our method demonstrates robust performance across varying context lengths, making it particularly suitable for diverse real-world applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang 等CVPR 2022 · 被引用 635 次
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang 等ICML 2022 · 被引用 523 次
- Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of ExpertsBasil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton 等NeurIPS 2022 · 被引用 359 次
相关 Paper
- CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet UpcyclingJihai Zhang, Xiaoye Qu, Tong Zhu, Yu ChengEMNLP 2025 · 被引用 1 次
- MoDE: CLIP Data Experts via ClusteringJiawei Ma, Po-Yao Huang, Saining Xie, Shang-Wen Li 等CVPR 2024 · 被引用 8 次
- Meta-Adapter: An Online Few-shot Learner for Vision-Language ModelCheng Cheng, Lin Song, Ruoyi Xue, Hang Wang 等NeurIPS 2023 · 被引用 65 次
- AmorLIP: Efficient Language-Image Pretraining via AmortizationHaotian Sun, Yitong Li, Yuchen Zhuang, Niao He 等NeurIPS 2025 · 被引用 2 次
- MoECLIP: Patch-Specialized Experts for Zero-shot Anomaly DetectionJun Yeong Park, JunYoung Seo, Minji Kang, Yu Rang ParkCVPR 2026
