Learning Granularity-Unified Representations for Text-to-Image Person Re-identification
Zhiyin Shao, Xinyu Zhang, Meng Fang, Zhifeng Lin, Jian Wang, Changxing Ding
Abstract
Text-to-image person re-identification (ReID) aims to search for pedestrian images of an interested identity via textual descriptions. It is challenging due to both rich intra-modal variations and significant inter-modal gaps. Existing works usually ignore the difference in feature granularity between the two modalities, i.e., the visual features are usually fine-grained while textual features are coarse, which is mainly responsible for the large inter-modal gaps. In this paper, we propose an end-to-end framework based on transformers to learn granularity-unified representations for both modalities, denoted as LGUR.
LGUR framework contains two modules: a Dictionary-based Granularity Alignment (DGA) module and a Prototype-based Granularity Unification (PGU) module. In DGA, in order to align the granularities of two modalities, we introduce a Multi-modality Shared Dictionary (MSD) to reconstruct both visual and textual features. Besides, DGA has two important factors, i.e., the cross-modality guidance and the foreground-centric reconstruction, to facilitate the optimization of MSD. In PGU, we adopt a set of shared and learnable prototypes as the queries to extract diverse and semantically aligned features for both modalities in the granularity-unified feature space, which further promotes the ReID performance. Comprehensive experiments show that our LGUR consistently outperforms state-of-the-arts by large margins on both CUHK-PEDES and ICFG-PEDES datasets. Code will be released at https://github.com/ZhiyinShao-H/LGUR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 187ac3d8-2241-48af-810d-a07fcf91050dCited by top-tier papers48
- Towards Unified Text-based Person Retrieval: A Large-scale Multi-Attribute and Language Search BenchmarkShuyu Yang, Yinan Zhou, Zhedong Zheng, Yaxiong Wang et al.ACM MM 2023 · 162 citations
- An Empirical Study of CLIP for Text-Based Person SearchMin Cao, Yang Bai, Ziyin Zeng, Mang Ye et al.AAAI 2024 · 111 citations
- PLIP: Language-Image Pre-training for Person Representation LearningJialong Zuo, Jiahao Hong, Feng Zhang, Changqian Yu et al.NeurIPS 2024 · 96 citations
- Noisy-Correspondence Learning for Text-to-Image Person Re-IdentificationYang Qin, Yingke Chen, Dezhong Peng, Xi Peng et al.CVPR 2024 · 83 citations
- Adaptive Uncertainty-Based Learning for Text-Based Person RetrievalShenshen Li, Chen He, Xing Xu, Fumin Shen et al.AAAI 2024 · 59 citations
Builds on12
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalAichun Zhu, Zijie Wang, Yifeng Li, Xili Wan et al.ACM MM 2021 · 274 citations
Related papers
- Unifying Multi-Modal Uncertainty Modeling and Semantic Alignment for Text-to-Image Person Re-identificationZhiwei Zhao, Bin Liu, Yan Lu, Qi Chu et al.AAAI 2024 · 40 citations
- Hierarchical Prompt Learning for Image- and Text-Based Person Re-IdentificationLinhan Zhou, Shuang Li, Neng Dong, Yonghang Tai et al.AAAI 2026 · 4 citations
- Text-Based Occluded Person Re-identification via Multi-Granularity Contrastive Consistency LearningXinyi Wu, Wentao Ma, Dan Guo, Tongqing Zhou et al.AAAI 2024 · 29 citations
- Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReIDWentao Tan, Changxing Ding, Jiayu Jiang, Fei Wang et al.CVPR 2024 · 34 citations
- Towards Grand Unified Representation Learning for Unsupervised Visible-Infrared Person Re-IdentificationBin Yang, Jun Chen, Mang YeICCV 2023 · 53 citations
