Probabilistic Embeddings for Cross-Modal Retrieval
Sanghyuk Chun, Seong Joon Oh, Rafael Sampaio de Rezende, Yannis Kalantidis, Diane Larlus
Abstract
Cross-modal retrieval methods build a common representation space for samples from multiple modalities, typically from the vision and the language domains. For images and their captions, the multiplicity of the correspondences makes the task particularly challenging. Given an image (respectively a caption), there are multiple captions (respectively images) that equally make sense. In this paper, we argue that deterministic functions are not sufficiently powerful to capture such one-to-many correspondences. Instead, we propose to use Probabilistic Cross-Modal Embedding (PCME), where samples from the different modalities are represented as probabilistic distributions in the common embedding space. Since common benchmarks such as COCO suffer from non-exhaustive annotations for crossmodal matches, we propose to additionally evaluate retrieval on the CUB dataset, a smaller yet clean database where all possible image-caption pairs are annotated. We extensively ablate PCME and demonstrate that it not only improves the retrieval performance over its deterministic counterpart but also provides uncertainty estimates that render the embeddings more interpretable. Code is available at https://github.com/naver-ai/pcme .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers87
- Learning with Noisy Correspondence for Cross-modal MatchingZhenyu Huang, Guocheng Niu, Xiao Liu, Wenbiao Ding et al.NeurIPS 2021 · 215 citations
- ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit SimilarityGinger Delmas, Rafael Sampaio de Rezende, Gabriela Csurka, Diane LarlusICLR 2022 · 147 citations
- Probabilistic Concept Bottleneck ModelsEunji Kim, Dahuin Jung, Sangha Park, Siwon Kim et al.ICML 2023 · 108 citations
- Expectation-Maximization Contrastive Learning for Compact Video-and-Language RepresentationsPeng Jin, Jinfa Huang, Fenglin Liu, Xian Wu et al.NeurIPS 2022 · 105 citations
- Exploring Diverse In-Context Configurations for Image CaptioningXu Yang, Yongliang Wu, Mingzhuo Yang, Haokun Chen et al.NeurIPS 2023 · 104 citations
Builds on11
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- Probabilistic Face EmbeddingsYichun Shi, Anil K. JainICCV 2019 · 362 citations
- CAMP: Cross-Modal Adaptive Message Passing for Text-Image RetrievalZihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng et al.ICCV 2019 · 349 citations
Related papers
- Improved Probabilistic Image-Text RepresentationsSanghyuk ChunICLR 2024 · 48 citations
- Deep Probabilistic Binary Embedding via Learning Reliable Uncertainty for Cross-Modal RetrievalKun Cheng, Qibing Qin, Wenfeng Zhang, Lei Huang et al.ACM MM 2025 · 4 citations
- A Differentiable Semantic Metric Approximation in Probabilistic Embedding for Cross-Modal RetrievalHao Li, Jingkuan Song, Lianli Gao, Pengpeng Zeng et al.NeurIPS 2022 · 22 citations
- ProbVLM: Probabilistic Adapter for Frozen Vison-Language ModelsUddeshya Upadhyay, Shyamgopal Karthik, Massimiliano Mancini, Zeynep AkataICCV 2023 · 41 citations
- Improving Cross-Modal Retrieval with Set of Diverse EmbeddingsDongwon Kim, Namyup Kim, Suha KwakCVPR 2023
