I0T: Embedding Standardization Method Towards Zero Modality Gap
Na Min An, Eunki Kim, James Thorne, Hyunjung Shim
Abstract
Contrastive Language-Image Pretraining (CLIP) enables zero-shot inference in downstream tasks such as image-text retrieval and classification. However, recent works extending CLIP suffer from the issue of modality gap, which arises when the image and text embeddings are projected to disparate manifolds, deviating from the intended objective of image-text contrastive learning. We discover that this phenomenon is linked to the modality-specific characteristic that each image or text encoder independently possesses. Herein, we propose two methods to address the modality gap: (1) a post-hoc embedding standardization method, I0T post that reduces the modality gap approximately to zero and (2) a trainable method, I0T async , to alleviate the modality gap problem by adding two normalization layers for each encoder. Our I0T framework can significantly reduce the modality gap while preserving the original embedding representations of trained models with their locked parameters. In practice, I0T post can serve as an alternative explainable automatic evaluation metric of widely used CLIPScore (CLIP-S). The code is available in https://github.com/xfactlab/I0T .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54d98bb4-722d-4611-bacc-ca109c3abeaeCited by top-tier papers4
- TextME: Bridging Unseen Modalities Through Text DescriptionsSoyeon Hong, Jinchan Kim, Jaegook You, Seungtaek Choi et al.ICML 2026
- Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via LanguageNam Hyeon-Woo, Yebin Moon, Sohwi Lim, Kwon Byung-Ki et al.ICML 2026
- Structured Multi-modal Graph Disentanglement for Psychiatric DiagnosisHongyu Shi, Kaizhong Zheng, WS Zhai, Shuai Jiang et al.ICML 2026
- Mitigating the Modality Gap in Vision–Language Models with Fractal Spectral GeometryZihan Zhou, Yang Zhou, Ruoming Jin, Pan He et al.ICML 2026
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Mitigate the Gap: Improving Cross-Modal Alignment in CLIPSedigheh Eslami, Gerard de MeloICLR 2025 · 1 citation
- Post-pre-training for Modality Alignment in Vision-Language Foundation ModelsShin'ya Yamaguchi, Dewei Feng, Sekitoshi Kanai, Kazuki Adachi et al.CVPR 2025
- Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality InversionMarco Mistretta, Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini et al.ICLR 2025
- Non-Contrastive Learning Meets Language-Image Pre-TrainingJinghao Zhou, Li Dong, Zhe Gan, Lijuan Wang et al.CVPR 2023
- Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual KnowledgeFawaz Sammani, Nikos DeligiannisNeurIPS 2024 · 15 citations
