FastPR: One-stage Semantic Person Retrieval via Self-supervised Learning
Meng Sun, Ju Ren, Xin Wang, Wenwu Zhu, Yaoxue Zhang
摘要
Semantic person retrieval aims to locate a specific person in an image with the query of semantic descriptions, which has shown great significance in surveillance and security applications. Prior arts commonly adopt a two-stage method that first extracts the persons with a pretrained detector and then finds the target matching the descriptions optimally.However, existing works suffer from high computational complexity and low recall rate caused by error accumulation in the two-stage inference. To solve the problems, we propose FastPR, a one-stage semantic person retrieval method via self-supervised learning, to optimize the person localization and semantic retrieval simultaneously. Specifically, we propose a dynamic visual-semantic alignment mechanism which utilizes grid-based attention to fuse the cross-modal features, and employs a label prediction proxy task to constrain the attention process. To tackle the challenges that real-world surveillance images may suffer from low-resolution and occlusion, and the target persons may be within a crowd,we further propose a dual-granularity person localization module through designing an upsampling reconstruction proxy task to enhance the local feature of the target person in the fused features, followed by a tailored offset prediction proxy task to make the localization network capable of accurately identifying and distinguishing the target person in a crowd. Experimental results demonstrate that FastPR achieves the best retrieval accuracy compared to the state-of-the-art baseline methods, with over 15 times inference time reduction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalAichun Zhu, Zijie Wang, Yifeng Li, Xili Wan 等ACM MM 2021 · 被引用 274 次
- Adversarial Representation Learning for Text-to-Image MatchingNikolaos Sarafianos, Xiang Xu, Ioannis A. KakadiarisICCV 2019 · 被引用 228 次
- Exploiting a Joint Embedding Space for Generalized Zero-Shot Semantic SegmentationDonghyeon Baek, Youngmin Oh, Bumsub HamICCV 2021 · 被引用 93 次
- Heterogeneous Attention Network for Effective and Efficient Cross-modal RetrievalTan Yu, Yi Yang, Yi Li, Lin Liu 等SIGIR 2021 · 被引用 50 次
相关 Paper
- Anchor-Free Person SearchYichao Yan, Jinpeng Li, Jie Qin, Song Bai 等CVPR 2021
- Self-similarity Driven Scale-invariant Learning for Weakly Supervised Person SearchBenzhi Wang, Yang Yang, Jinlin Wu, Guo-Jun Qi 等ICCV 2023 · 被引用 15 次
- Pedestrian-Centric Discriminative and Fine-grained Semantic Mining for Text-based Person RetrievalYuheng Liang, Haipeng Chen, Yu Liu, Yingda Lyu 等WWW 2026
- Cross-modal Co-occurrence Attributes Alignments for Person Search by LanguageKai Niu, Linjiang Huang, Yan Huang, Peng Wang 等ACM MM 2022 · 被引用 35 次
- Fine-grained Semantic Alignment with Transferred Person-SAM for Text-based Person RetrievalYihao Wang, Meng Yang, Rui CaoACM MM 2024 · 被引用 15 次
