Fine-grained Semantic Alignment with Transferred Person-SAM for Text-based Person Retrieval
Yihao Wang, Meng Yang, Rui Cao
Abstract
Addressing the disparity in description granularity and information gap between images and text has long been a formidable challenge in text-based person retrieval (TBPR) tasks. Recent researchers tried to solve this problem by random local alignment. However, they failed to capture the fine-grained relationships between images and text, so the information and modality gaps remain on the table. We align image regions and text phrases at the same semantic granularity to address the semantic atomicity gap. Our idea is first to extract and then exploit the relationships between fine-grained locals. We introduce a novel Fine-grained Semantic Alignment with Transferred Person-SAM (SAP-SAM) approach. By distilling and transferring knowledge, we propose a Person-SAM model to extract fine-grained semantic concepts at the same granularity from images and texts of TBPR and its relationships. With the extracted knowledge, we optimize the fine-grained matching via Explicit Local Concept Alignment and Attentive Cross-modal Decoding to discriminate fine-grained image and text features at the same granularity level and represent the important semantic concepts from both modalities, effectively alleviating the granularity and information gaps. We evaluate our proposed approach on three popular TBPR datasets, demonstrating that SAP-SAM achieves state-of-the-art results and underscores the effectiveness of end-to-end fine-grained local alignment in TBPR tasks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8a5f6642-4b2a-43ea-b09f-e2824e5ee3d3Cited by top-tier papers3
- CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ SegmentationXinlei Yu, Changmiao Wang, Hui Jin, Ahmed Elazab et al.ACM MM 2025 · 3 citations
- KPDM: Key Phrase Dynamic Masking for Robust Text-to-Image Person RetrievalShaofeng You, Tianle Miao, Qihang Chen, Xin Li et al.AAAI 2026
- CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image RetrievalBin Kang, Bin Chen, Junjie Wang, Yulin Li et al.ACM MM 2025
Related papers
- ASMR: Learning Attribute-Based Person Search with Adaptive Semantic Margin RegularizerBoseung Jeong, Jicheol Park, Suha KwakICCV 2021 · 29 citations
- Pedestrian-Centric Discriminative and Fine-grained Semantic Mining for Text-based Person RetrievalYuheng Liang, Haipeng Chen, Yu Liu, Yingda Lyu et al.WWW 2026
- Learning Hierarchical Cross-modal Association with Intra-modal Context for Text-Image Person RetrievalYifei Deng, Chenglong Li, Futian Wang, Jin TangACM MM 2025 · 2 citations
- UFineBench: Towards Text-based Person Retrieval with Ultra-fine GranularityJialong Zuo, Hanyu Zhou, Ying Nie, Feng Zhang et al.CVPR 2024 · 45 citations
- Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person RetrievalDing Jiang, Mang YeCVPR 2023
