Is CLIP Ideal? No. Can We Fix It? Yes!
Raphi Kang, Yue Song, Georgia Gkioxari, Pietro Perona
Abstract
Contrastive Language-Image Pre-Training (CLIP) is a popular method for learning multimodal latent spaces with well-organized semantics. Despite its wide range of applications, CLIP's latent space is known to fail at handling complex visual-textual interactions. Recent works attempt to address its shortcomings with data-centric or algorithmic approaches. But what if the problem is more fundamental, and lies in the geometry of CLIP? Toward this end, we rigorously analyze CLIP's latent space properties, and prove that no CLIP-like joint embedding space exists which can correctly do any two of the following at the same time: 1. represent basic descriptions and image content, 2. represent attribute binding, 3. represent spatial location and relationships, 4. represent negation. Informed by this analysis, we propose Dense Cosine Similarity Maps (DCSMs) as a principled and interpretable scoring method for CLIP-like models, which solves the fundamental limitations of CLIP by retaining the semantic topology of the image patches and text tokens. This method improves upon the performance of classical CLIP-like joint encoder models on a wide array of benchmarks. We share our code and data here for reproducibility: https: // github. com/ Raphoo/ DCSM_ Ideal_ CLIP
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Circuit Mechanisms for Spatial Relation Generation in Diffusion TransformersBinxu Wang, Jingxuan Fan, Xu PanCVPR 2026 · 4 citations
- PowerCLIP: Powerset Alignment for Contrastive Pre-TrainingMasaki Kawamura, Nakamasa Inoue, Rintaro Yanagi, Hirokatsu Kataoka et al.CVPR 2026 · 1 citation
- SimLBR: Learning to Detect Fake Images by Learning to Detect Real ImagesAayush Dhakal, Subash Khanal, Srikumar Sastry, Jacob Arndt et al.CVPR 2026 · 1 citation
- PhaseAlign: Complex Phase Alignment for Stable Open-Vocabulary Semantic SegmentationJiankang Wang, Dingding Jia, Zhoushuopeng, Xuan WangICML 2026
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modallyDarina Koishigarina, Arnas Uselis, Seong Joon OhICLR 2026 · 33 citations
- Mitigate the Gap: Improving Cross-Modal Alignment in CLIPSedigheh Eslami, Gerard de MeloICLR 2025 · 1 citation
- Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic SegmentationXianglin Qiu, Jian Wang, Xiaolei Wang, Zhen Zhang et al.CVPR 2026
- Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual KnowledgeFawaz Sammani, Nikos DeligiannisNeurIPS 2024 · 15 citations
- un2CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIPYinqi Li, Jiahe Zhao, Hong Chang, Ruibing Hou et al.NeurIPS 2025 · 6 citations
