Pose-Guided Multi-Granularity Attention Network for Text-Based Person Search
Ya Jing, Chenyang Si, Junbo Wang, Wei Wang, Liang Wang, Tieniu Tan
Abstract
Text-based person search aims to retrieve the corresponding person images in an image database by virtue of a describing sentence about the person, which poses great potential for various applications such as video surveillance. Extracting visual contents corresponding to the human description is the key to this cross-modal matching problem. Moreover, correlated images and descriptions involve different granularities of semantic relevance, which is usually ignored in previous methods. To exploit the multilevel corresponding visual contents, we propose a pose-guided multi-granularity attention network (PMA). Firstly, we propose a coarse alignment network (CA) to select the related image regions to the global description by a similarity-based attention. To further capture the phrase-related visual body part, a fine-grained alignment network (FA) is proposed, which employs pose information to learn latent semantic alignment between visual body part and textual noun phrase. To verify the effectiveness of our model, we perform extensive experiments on the CUHK Person Description Dataset (CUHK-PEDES) which is currently the only available dataset for text-based person search. Experimental results show that our approach outperforms the state-of-the-art methods by 15 % in terms of the top-1 metric.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d91467c3-b30f-4f51-9728-1edae0cdba7fCited by top-tier papers30
- DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalAichun Zhu, Zijie Wang, Yifeng Li, Xili Wan et al.ACM MM 2021 · 274 citations
- Learning Granularity-Unified Representations for Text-to-Image Person Re-identificationZhiyin Shao, Xinyu Zhang, Meng Fang, Zhifeng Lin et al.ACM MM 2022 · 197 citations
- Towards Unified Text-based Person Retrieval: A Large-scale Multi-Attribute and Language Search BenchmarkShuyu Yang, Yinan Zhou, Zhedong Zheng, Yaxiong Wang et al.ACM MM 2023 · 162 citations
- AXM-Net: Implicit Cross-Modal Feature Alignment for Person Re-identificationAmmarah Farooq, Muhammad Awais, Josef Kittler, Syed Safwan KhalidAAAI 2022 · 126 citations
- CAIBC: Capturing All-round Information Beyond Color for Text-based Person RetrievalZijie Wang, Aichun Zhu, Jingyi Xue, Xili Wan et al.ACM MM 2022 · 126 citations
Related papers
- Cross-Modal Cross-Domain Moment Alignment Network for Person SearchYa Jing, Wei Wang, Liang Wang, Tieniu TanCVPR 2020
- Hierarchical Gumbel Attention Network for Text-based Person SearchKecheng Zheng, Wu Liu, Jiawei Liu, Zheng-Jun Zha et al.ACM MM 2020 · 96 citations
- Cross-modal Co-occurrence Attributes Alignments for Person Search by LanguageKai Niu, Linjiang Huang, Yan Huang, Peng Wang et al.ACM MM 2022 · 35 citations
- Textual Dependency Embedding for Person Search by LanguageKai Niu, Yan Huang, Liang WangACM MM 2020 · 42 citations
- Multi-Modal Disordered Representation Learning Network for Description-Based Person SearchFan Yang, Wei Li, Menglong Yang, Binbin Liang et al.AAAI 2024 · 8 citations
