Probabilistic Embeddings for Cross-Modal Retrieval
Sanghyuk Chun, Seong Joon Oh, Rafael Sampaio de Rezende, Yannis Kalantidis, Diane Larlus
摘要
Cross-modal retrieval methods build a common representation space for samples from multiple modalities, typically from the vision and the language domains. For images and their captions, the multiplicity of the correspondences makes the task particularly challenging. Given an image (respectively a caption), there are multiple captions (respectively images) that equally make sense. In this paper, we argue that deterministic functions are not sufficiently powerful to capture such one-to-many correspondences. Instead, we propose to use Probabilistic Cross-Modal Embedding (PCME), where samples from the different modalities are represented as probabilistic distributions in the common embedding space. Since common benchmarks such as COCO suffer from non-exhaustive annotations for crossmodal matches, we propose to additionally evaluate retrieval on the CUB dataset, a smaller yet clean database where all possible image-caption pairs are annotated. We extensively ablate PCME and demonstrate that it not only improves the retrieval performance over its deterministic counterpart but also provides uncertainty estimates that render the embeddings more interpretable. Code is available at https://github.com/naver-ai/pcme .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper87
- Learning with Noisy Correspondence for Cross-modal MatchingZhenyu Huang, Guocheng Niu, Xiao Liu, Wenbiao Ding 等NeurIPS 2021 · 被引用 215 次
- ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit SimilarityGinger Delmas, Rafael Sampaio de Rezende, Gabriela Csurka, Diane LarlusICLR 2022 · 被引用 147 次
- Probabilistic Concept Bottleneck ModelsEunji Kim, Dahuin Jung, Sangha Park, Siwon Kim 等ICML 2023 · 被引用 108 次
- Expectation-Maximization Contrastive Learning for Compact Video-and-Language RepresentationsPeng Jin, Jinfa Huang, Fenglin Liu, Xian Wu 等NeurIPS 2022 · 被引用 105 次
- Exploring Diverse In-Context Configurations for Image CaptioningXu Yang, Yongliang Wu, Mingzhuo Yang, Haokun Chen 等NeurIPS 2023 · 被引用 104 次
它引用的顶会 Paper11
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li 等ICCV 2019 · 被引用 598 次
- Probabilistic Face EmbeddingsYichun Shi, Anil K. JainICCV 2019 · 被引用 362 次
- CAMP: Cross-Modal Adaptive Message Passing for Text-Image RetrievalZihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng 等ICCV 2019 · 被引用 349 次
相关 Paper
- Improved Probabilistic Image-Text RepresentationsSanghyuk ChunICLR 2024 · 被引用 48 次
- Deep Probabilistic Binary Embedding via Learning Reliable Uncertainty for Cross-Modal RetrievalKun Cheng, Qibing Qin, Wenfeng Zhang, Lei Huang 等ACM MM 2025 · 被引用 4 次
- A Differentiable Semantic Metric Approximation in Probabilistic Embedding for Cross-Modal RetrievalHao Li, Jingkuan Song, Lianli Gao, Pengpeng Zeng 等NeurIPS 2022 · 被引用 22 次
- ProbVLM: Probabilistic Adapter for Frozen Vison-Language ModelsUddeshya Upadhyay, Shyamgopal Karthik, Massimiliano Mancini, Zeynep AkataICCV 2023 · 被引用 41 次
- Improving Cross-Modal Retrieval with Set of Diverse EmbeddingsDongwon Kim, Namyup Kim, Suha KwakCVPR 2023
