A Differentiable Semantic Metric Approximation in Probabilistic Embedding for Cross-Modal Retrieval
Hao Li, Jingkuan Song, Lianli Gao, Pengpeng Zeng, Haonan Zhang, Gongfu Li
摘要
Cross-modal retrieval aims to build correspondence between multiple modalities by learning a common representation space. Typically, an image can match multiple texts semantically and vice versa, which significantly increases the difficulty of this task. To address this problem, probabilistic embedding is proposed to quantify these many-to-many relationships. However, existing datasets ( e . g ., MS-COCO) and metrics ( e . g ., Recall@K) cannot fully represent these diversity correspondences due to non-exhaustive annotations. Based on this observation, we utilize semantic correlation computed by CIDEr to find the potential correspondences. Then we present an effective metric, named Average Semantic Precision (ASP), which can measure the ranking precision of semantic correlation for retrieval sets. Additionally, we introduce a novel and concise objective, coined Differentiable ASP Approximation (DAA). Concretely, DAA can optimize ASP directly by making the ranking function of ASP differentiable through a sigmoid function. To verify the effectiveness of our approach, extensive experiments are conducted on MS-COCO, CUB Captions, and Flickr30K, which are commonly used in cross-modal retrieval. The results show that our approach obtains superior performance over the state-of-the-art approaches on all metrics. The code and trained models are released at https://github.com/leolee99/2022-NeurIPS-DAA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Prototype-based Aleatoric Uncertainty Quantification for Cross-modal RetrievalHao Li, Jingkuan Song, Lianli Gao, Xiaosu Zhu 等NeurIPS 2023 · 被引用 52 次
- Improved Probabilistic Image-Text RepresentationsSanghyuk ChunICLR 2024 · 被引用 48 次
- ProbVLM: Probabilistic Adapter for Frozen Vison-Language ModelsUddeshya Upadhyay, Shyamgopal Karthik, Massimiliano Mancini, Zeynep AkataICCV 2023 · 被引用 41 次
- Negative Pre-aware for Noisy Cross-Modal MatchingXu Zhang, Hao Li, Mang YeAAAI 2024 · 被引用 18 次
- Post-hoc Probabilistic Vision-Language ModelsAnton Baumann, Rui Li, Marcus Klasson, Santeri Mentu 等ICLR 2026 · 被引用 14 次
它引用的顶会 Paper14
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li 等ICCV 2019 · 被引用 598 次
- Learning With Average Precision: Training Image Retrieval With a Listwise LossJérôme Revaud, Jon Almazán, Rafael S. Rezende, César Roberto de SouzaICCV 2019 · 被引用 424 次
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 被引用 413 次
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant WeightsByeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han 等ICLR 2021 · 被引用 165 次
- SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic QuantizationYuhta Takida, Takashi Shibuya, Wei-Hsiang Liao, Chieh-Hsin Lai 等ICML 2022 · 被引用 99 次
相关 Paper
- Probabilistic Embeddings for Cross-Modal RetrievalSanghyuk Chun, Seong Joon Oh, Rafael Sampaio de Rezende, Yannis Kalantidis 等CVPR 2021
- Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text MatchingXiang Ma, Xuemei Li, Lexin Fang, Caiming ZhangACM MM 2024 · 被引用 4 次
- Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal RetrievalHani Alomari, Anushka Sivakumar, Andrew Zhang, Chris ThomasACL 2025
- Learning Relation Alignment for Calibrated Cross-modal RetrievalShuhuai Ren, Junyang Lin, Guangxiang Zhao, Rui Men 等ACL 2021
- Adaptive Prompt-Based Semantic Embedding with Inspire Potential of Implicit Knowledge for Cross-Modal RetrievalXin Huang, Shilong Wang, Tong Jia, Zhihang Gou 等AAAI 2025 · 被引用 2 次
