Point to Rectangle Matching for Image Text Retrieval
Zheng Wang, Zhenwei Gao, Xing Xu, Yadan Luo, Yang Yang, Heng Tao Shen
Abstract
The difficulty of image-text retrieval is further exacerbated by the phenomenon of one-to-many correspondence, where multiple semantic manifestations of the other modality could be obtained by a given query. However, the prevailing methods adopt the deterministic embedding strategy to retrieve the most similar candidate, which encodes the representations of different modalities as single points in vector space. We argue that such a deterministic point mapping is obviously insufficient to represent a potential set of retrieval results for one-to-many correspondence, despite its noticeable progress. As a remedy to this issue, we propose a Point to Rectangle Matching (abbreviated as P2RM) mechanism, which actually is a geometric representation learning method for image-text retrieval. Specifically, our intuitive insight is that the representations of different modalities could be extended to rectangles, then a set of points inside such a rectangle embedding could be semantically related to many candidate correspondences. Thus our P2RM method could essentially address the one-to-many correspondence. Besides, we design a novel semantic similarity measurement method from the perspective of distance for our rectangle embedding. Under the evaluation metric for multiple matches, extensive experiments and ablation studies on two commonly used benchmarks demonstrate our effectiveness and superiority in tackling the multiplicity of image-text retrieval.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ea975667-442c-41b8-987a-70fde7c97278Cited by top-tier papers6
- Improved Probabilistic Image-Text RepresentationsSanghyuk ChunICLR 2024 · 48 citations
- Your Negative May not Be True Negative: Boosting Image-Text Matching with False Negative EliminationHaoxuan Li, Yi Bin, Junrong Liao, Yang Yang et al.ACM MM 2023 · 42 citations
- Open-Vocabulary Object Detection via Scene Graph DiscoveryHengcan Shi, Munawar Hayat, Jianfei CaiACM MM 2023 · 20 citations
- CDPNet: Cross-Modal Dual Phases Network for Point Cloud CompletionZhenjiang Du, Jiale Dou, Zhitao Liu, Jiwei Wei et al.AAAI 2024 · 17 citations
- Variational Adapter for Cross-modal Similarity RepresentationWenZhang Wei, Zhipeng Gui, Dehua Peng, Tiandi Ye et al.ICML 2026 · 1 citation
Related papers
- Multilateral Semantic Relations Modeling for Image Text RetrievalZheng Wang, Zhenwei Gao, Kangshuai Guo, Yang Yang et al.CVPR 2023
- Probabilistic Embeddings for Cross-Modal RetrievalSanghyuk Chun, Seong Joon Oh, Rafael Sampaio de Rezende, Yannis Kalantidis et al.CVPR 2021
- Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal RetrievalHani Alomari, Anushka Sivakumar, Andrew Zhang, Chris ThomasACL 2025
- HL-CMR: Hypergraph Learning for Cross-Modal RetrievalGuohui Ding, Jing Li, Yimin Xu, Rui ZhouWWW 2026
- Text-Adaptive Multiple Visual Prototype Matching for Video-Text RetrievalChengzhi Lin, Ancong Wu, Junwei Liang, Jun Zhang et al.NeurIPS 2022 · 52 citations
