MLIP: Efficient Multi-Perspective Language-Image Pretraining with Exhaustive Data Utilization
Yu Zhang, Qi Zhang, Zixuan Gong, Yiwei Shi, Yepeng Liu, Duoqian Miao, Yang Liu, Ke Liu, Kun Yi, Wei Fan, Liang Hu, Changwei Wang
Abstract
Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success, leading to rapid advancements in multimodal studies. However, CLIP faces a notable challenge in terms of inefficient data utilization. It relies on a single contrastive supervision for each image-text pair during representation learning, disregarding a substantial amount of valuable information that could offer richer supervision. Additionally, the retention of non-informative tokens leads to increased computational demands and time costs, particularly in CLIP's ViT image encoder. To address these issues, we propose Multi-Perspective Language-Image Pretraining (MLIP). In MLIP, we leverage the frequency transform's sensitivity to both high and low-frequency variations, which complements the spatial domain's sensitivity limited to low-frequency variations only. By incorporating frequency transforms and token-level alignment, we expand CILP's single supervision into multi-domain and multi-level supervision, enabling a more thorough exploration of informative image features. Additionally, we introduce a token merging method guided by comprehensive semantics from the frequency and spatial domains. This allows us to merge tokens to multi-granularity tokens with a controllable compression rate to accelerate CLIP. Extensive experiments validate the effectiveness of our design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 464f6d41-db81-430e-9a56-32c3071a4a3cCited by top-tier papers3
- NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video ReconstructionZixuan Gong, Guangyin Bao, Qi Zhang, Zhongwei Wan et al.NeurIPS 2024 · 39 citations
- Enhancing Text-to-Image Diffusion Transformer via Split-Text ConditioningYu Zhang, Jialei Zhou, Xinchen Li, Qi Zhang et al.NeurIPS 2025 · 11 citations
- Lite-Mind: Towards Efficient and Robust Brain Representation LearningZixuan Gong, Qi Zhang, Guangyin Bao, Lei Zhu et al.ACM MM 2024 · 2 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FcaNet: Frequency Channel Attention NetworksZequn Qin, Pengyi Zhang, Fei Wu, Xi LiICCV 2021 · 1,049 citations
- Spectral Temporal Graph Neural Network for Multivariate Time-series ForecastingDefu Cao, Yujing Wang, Juanyong Duan, Ce Zhang et al.NeurIPS 2020 · 841 citations
- Global Filter Networks for Image ClassificationYongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu et al.NeurIPS 2021 · 798 citations
Related papers
- UniCLIP: Unified Framework for Contrastive Language-Image Pre-trainingJanghyeon Lee, Jongsuk Kim, Hyounguk Shon, Bumsoo Kim et al.NeurIPS 2022 · 85 citations
- OneLIP: Unlocking and Improving Long-Text Representations of CLIP via One-Stage AdaptationRenjie Pan, Jiayan Song, Hua YangAAAI 2026
- Non-Contrastive Learning Meets Language-Image Pre-TrainingJinghao Zhou, Li Dong, Zhe Gan, Lijuan Wang et al.CVPR 2023
- FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained UnderstandingDong Jing, Xiaolong He, Yutian Luo, Nanyi Fei et al.NeurIPS 2024 · 70 citations
- MLIP: Enhancing Medical Visual Representation with Divergence Encoder and Knowledge-guided Contrastive LearningZhe Li, Laurence T. Yang, Bocheng Ren, Xin Nie et al.CVPR 2024
