Learnable Pillar-based Re-ranking for Image-Text Retrieval
Leigang Qu, Meng Liu, Wenjie Wang, Zhedong Zheng, Liqiang Nie, Tat-Seng Chua
Abstract
Image-text retrieval aims to bridge the modality gap and retrieve cross-modal content based on semantic similarities. Prior work usually focuses on the pairwise relations (i.e., whether a data sample matches another) but ignores the higher-order neighbor relations (i.e., a matching structure among multiple data samples). Re-ranking, a popular post-processing practice, has revealed the superiority of capturing neighbor relations in single-modality retrieval tasks. However, it is ineffective to directly extend existing re-ranking algorithms to image-text retrieval. In this paper, we analyze the reason from four perspectives, i.e., generalization, flexibility, sparsity, and asymmetry, and propose a novel learnable pillar-based re-ranking paradigm. Concretely, we first select top-ranked intra- and intermodal neighbors as pillars, and then reconstruct data samples with the neighbor relations between them and the pillars. In this way, each sample can be mapped into a multimodal pillar space only using similarities, ensuring generalization. After that, we design a neighbor-aware graph reasoning module to flexibly exploit the relations and excavate the sparse positive items within a neighborhood. We also present a structure alignment constraint to promote crossmodal collaboration and align the asymmetric modalities. On top of various base backbones, we carry out extensive experiments on two benchmark datasets, i.e., Flickr30K and MS-COCO, demonstrating the effectiveness, superiority, generalization, and transferability of our proposed re-ranking paradigm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei et al.ACM MM 2023 · 53 citations
- CaLa: Complementary Association Learning for Augmenting Comoposed Image RetrievalXintong Jiang, Yaxiong Wang, Mengjian Li, Yujiao Wu et al.SIGIR 2024 · 13 citations
- CFIR: Fast and Effective Long-Text To Image Retrieval for Large CorporaZijun Long, Xuri Ge, Richard McCreadie, Joemon M. JoseSIGIR 2024 · 10 citations
- Multimodal Hypothetical Summary for Retrieval-based Multi-image Question AnsweringPeize Li, Qingyi Si, Peng Fu, Zheng Lin et al.AAAI 2025 · 1 citation
- CarGait: Cross-Attention Based Re-ranking for Gait RecognitionGavriel Habib, Noa Barzilay, Or Shimshi, Rami Ben-Ari et al.ICCV 2025
Builds on16
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- Learning With Average Precision: Training Image Retrieval With a Listwise LossJérôme Revaud, Jon Almazán, Rafael S. Rezende, César Roberto de SouzaICCV 2019 · 424 citations
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
- Dynamic Modality Interaction Modeling for Image-Text RetrievalLeigang Qu, Meng Liu, Jianlong Wu, Zan Gao et al.SIGIR 2021 · 187 citations
- Context-Aware Multi-View Summarization Network for Image-Text MatchingLeigang Qu, Meng Liu, Da Cao, Liqiang Nie et al.ACM MM 2020 · 159 citations
Related papers
- Learning Semantic Relationship among Instances for Image-Text MatchingZheren Fu, Zhendong Mao, Yan Song, Yongdong ZhangCVPR 2023
- HL-CMR: Hypergraph Learning for Cross-Modal RetrievalGuohui Ding, Jing Li, Yimin Xu, Rui ZhouWWW 2026
- Intra-Modal Neighbors Never Lie: Rectifying Inter-Modal Noisy Correspondence via Graph-Based Intra-Modal ReasoningYang Liu, Wentao Feng, Shudong Huang, Yalan Ye et al.ICML 2026
- Exploring Graph-Structured Semantics for Cross-Modal RetrievalLei Zhang, Leiting Chen, Chuan Zhou, Fan Yang et al.ACM MM 2021 · 14 citations
- Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person RetrievalDing Jiang, Mang YeCVPR 2023
