OT-CLIP: Understanding and Generalizing CLIP via Optimal Transport
Liangliang Shi, Jack Fan, Junchi Yan
摘要
We propose to understand Contrastive Language-Image Pretraining model (CLIP) from the Optimal Transport (OT) perspective. Specifically, we show that training of CLIP is an embodiment of inverse OT and the adopted two InfoNCE losses in CLIP correspond to a special case of bilevel optimization of a modified entropic OT. We then generalize the original CLIP loss to an OT-based loss family using variants of Regularized OT (e.g. Fused Gromov OT, unbalanced OT, etc.), and show their superior performance on public datasets for downstream tasks in both image and text domain. We also rethink the inference stage of CLIP by using the tool of OT, and propose to adopt the fused Gromov OT for (zero-shot) classification, in which the prediction is based on the graph representation whereby images and texts are nodes for graph matching. By our new technique, we show how to generalize zero-shot classification to other more flexible zero-shot tasks with competitive performance: long-tailed classification and selective classification. The former assumes the known prior distribution of labels, while in the latter case, only a subset of samples are asked to predict, yet with the need of high prediction confidence. The code is available at https://github.com/fan23j/ICML2024-OT-CLIP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- KPL: Training-Free Medical Knowledge Mining of Vision-Language ModelsJiaxiang Liu, Tianxiang Hu, Jiawei Du, Ruiyuan Zhang 等AAAI 2025 · 被引用 9 次
- Computing Approximate Graph Edit Distance via Optimal TransportQihao Cheng, Da Yan, Tianhao Wu, Zhongyi Huang 等SIGMOD 2025 · 被引用 5 次
- DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion ModelsQichao Wang, Yunhong Lu, Hengyuan Cao, Junyi Zhang 等CVPR 2026 · 被引用 4 次
- Unlearning the Noisy Correspondence Makes CLIP More RobustHaochen Han, Alex Jinpeng Wang, Peijun Ye, Fangming LiuICCV 2025 · 被引用 3 次
- PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned RewardsMinh-Quan Le, Gaurav Mittal, Cheng Zhao, Xianfeng GU 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
相关 Paper
- Data Efficient Language-Supervised Zero-Shot Recognition with Optimal Transport DistillationBichen Wu, Ruizhe Cheng, Peizhao Zhang, Tianren Gao 等ICLR 2022 · 被引用 57 次
- Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and AlignmentTong Zhang, Kuofeng Gao, Jiawang Bai, Leo Yu Zhang 等EMNLP 2025 · 被引用 1 次
- Inverse Optimal Transport for Efficient Adaptation of Vision-Language ModelsShupeng Qiu, Chuan-Xian RenAAAI 2026 · 被引用 1 次
- Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual KnowledgeFawaz Sammani, Nikos DeligiannisNeurIPS 2024 · 被引用 15 次
- CLIPood: Generalizing CLIP to Out-of-DistributionsYang Shu, Xingzhuo Guo, Jialong Wu, Ximei Wang 等ICML 2023 · 被引用 122 次
