FastPR: One-stage Semantic Person Retrieval via Self-supervised Learning
Meng Sun, Ju Ren, Xin Wang, Wenwu Zhu, Yaoxue Zhang
Abstract
Semantic person retrieval aims to locate a specific person in an image with the query of semantic descriptions, which has shown great significance in surveillance and security applications. Prior arts commonly adopt a two-stage method that first extracts the persons with a pretrained detector and then finds the target matching the descriptions optimally.However, existing works suffer from high computational complexity and low recall rate caused by error accumulation in the two-stage inference. To solve the problems, we propose FastPR, a one-stage semantic person retrieval method via self-supervised learning, to optimize the person localization and semantic retrieval simultaneously. Specifically, we propose a dynamic visual-semantic alignment mechanism which utilizes grid-based attention to fuse the cross-modal features, and employs a label prediction proxy task to constrain the attention process. To tackle the challenges that real-world surveillance images may suffer from low-resolution and occlusion, and the target persons may be within a crowd,we further propose a dual-granularity person localization module through designing an upsampling reconstruction proxy task to enhance the local feature of the target person in the fused features, followed by a tailored offset prediction proxy task to make the localization network capable of accurately identifying and distinguishing the target person in a crowd. Experimental results demonstrate that FastPR achieves the best retrieval accuracy compared to the state-of-the-art baseline methods, with over 15 times inference time reduction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84fe05ff-4330-4461-806a-076eefa1ad5aBuilds on4
- DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalAichun Zhu, Zijie Wang, Yifeng Li, Xili Wan et al.ACM MM 2021 · 274 citations
- Adversarial Representation Learning for Text-to-Image MatchingNikolaos Sarafianos, Xiang Xu, Ioannis A. KakadiarisICCV 2019 · 228 citations
- Exploiting a Joint Embedding Space for Generalized Zero-Shot Semantic SegmentationDonghyeon Baek, Youngmin Oh, Bumsub HamICCV 2021 · 93 citations
- Heterogeneous Attention Network for Effective and Efficient Cross-modal RetrievalTan Yu, Yi Yang, Yi Li, Lin Liu et al.SIGIR 2021 · 50 citations
Related papers
- Anchor-Free Person SearchYichao Yan, Jinpeng Li, Jie Qin, Song Bai et al.CVPR 2021
- Self-similarity Driven Scale-invariant Learning for Weakly Supervised Person SearchBenzhi Wang, Yang Yang, Jinlin Wu, Guo-Jun Qi et al.ICCV 2023 · 15 citations
- Pedestrian-Centric Discriminative and Fine-grained Semantic Mining for Text-based Person RetrievalYuheng Liang, Haipeng Chen, Yu Liu, Yingda Lyu et al.WWW 2026
- Cross-modal Co-occurrence Attributes Alignments for Person Search by LanguageKai Niu, Linjiang Huang, Yan Huang, Peng Wang et al.ACM MM 2022 · 35 citations
- Fine-grained Semantic Alignment with Transferred Person-SAM for Text-based Person RetrievalYihao Wang, Meng Yang, Rui CaoACM MM 2024 · 15 citations
