OT-CLIP: Understanding and Generalizing CLIP via Optimal Transport
Liangliang Shi, Jack Fan, Junchi Yan
Abstract
We propose to understand Contrastive Language-Image Pretraining model (CLIP) from the Optimal Transport (OT) perspective. Specifically, we show that training of CLIP is an embodiment of inverse OT and the adopted two InfoNCE losses in CLIP correspond to a special case of bilevel optimization of a modified entropic OT. We then generalize the original CLIP loss to an OT-based loss family using variants of Regularized OT (e.g. Fused Gromov OT, unbalanced OT, etc.), and show their superior performance on public datasets for downstream tasks in both image and text domain. We also rethink the inference stage of CLIP by using the tool of OT, and propose to adopt the fused Gromov OT for (zero-shot) classification, in which the prediction is based on the graph representation whereby images and texts are nodes for graph matching. By our new technique, we show how to generalize zero-shot classification to other more flexible zero-shot tasks with competitive performance: long-tailed classification and selective classification. The former assumes the known prior distribution of labels, while in the latter case, only a subset of samples are asked to predict, yet with the need of high prediction confidence. The code is available at https://github.com/fan23j/ICML2024-OT-CLIP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32c8c0dc-8c24-4e83-9dab-990c58e40fadCited by top-tier papers14
- KPL: Training-Free Medical Knowledge Mining of Vision-Language ModelsJiaxiang Liu, Tianxiang Hu, Jiawei Du, Ruiyuan Zhang et al.AAAI 2025 · 9 citations
- Computing Approximate Graph Edit Distance via Optimal TransportQihao Cheng, Da Yan, Tianhao Wu, Zhongyi Huang et al.SIGMOD 2025 · 5 citations
- DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion ModelsQichao Wang, Yunhong Lu, Hengyuan Cao, Junyi Zhang et al.CVPR 2026 · 4 citations
- Unlearning the Noisy Correspondence Makes CLIP More RobustHaochen Han, Alex Jinpeng Wang, Peijun Ye, Fangming LiuICCV 2025 · 3 citations
- PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned RewardsMinh-Quan Le, Gaurav Mittal, Cheng Zhao, Xianfeng GU et al.ICML 2026 · 2 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
Related papers
- Data Efficient Language-Supervised Zero-Shot Recognition with Optimal Transport DistillationBichen Wu, Ruizhe Cheng, Peizhao Zhang, Tianren Gao et al.ICLR 2022 · 57 citations
- Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and AlignmentTong Zhang, Kuofeng Gao, Jiawang Bai, Leo Yu Zhang et al.EMNLP 2025 · 1 citation
- Inverse Optimal Transport for Efficient Adaptation of Vision-Language ModelsShupeng Qiu, Chuan-Xian RenAAAI 2026 · 1 citation
- Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual KnowledgeFawaz Sammani, Nikos DeligiannisNeurIPS 2024 · 15 citations
- CLIPood: Generalizing CLIP to Out-of-DistributionsYang Shu, Xingzhuo Guo, Jialong Wu, Ximei Wang et al.ICML 2023 · 122 citations
