A Pedestrian is Worth One Prompt: Towards Language Guidance Person Re- Identification
Zexian Yang, Dayan Wu, Chenming Wu, Zheng Lin, Jingzi Gu, Weiping Wang
Abstract
Extensive advancements have been made in person ReID through the mining of semantic information. Nevertheless, existing methods that utilize semantic-parts from a single image modality do not explicitly achieve this goal. Whiteness the impressive capabilities in multimodal understanding of Vision Language Foundation Model CLIP, a recent two-stage CLIP-based method employs automated prompt engineering to obtain specific textual labels for classifying pedestrians. However, we note that the predefined soft prompts may be inadequate in expressing the entire visual context and struggle to generalize to unseen classes. This paper presents an end-to-end Prompt-driven Semantic Guidance (PromptSG) framework that harnesses the rich semantics inherent in CLIP. Specifically, we guide the model to attend to regions that are semantically faithful to the prompt. To provide personalized language descriptions for specific individuals, we propose learning pseudo tokens that represent specific visual contexts. This design not only facilitates learning fine-grained attribute information but also can inherently leverage language prompts during inference. Without requiring additional labeling efforts, our PromptSG achieves state-of-the-art by over 10% on MSMTI7 and nearly 5% on the Market-I50I benchmark. The codes will be available at h t tps: / / gi th ub. com/ YzXian16/PromptSG
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c706c48c-c1e7-4bf2-a0ae-f116ad8e196eCited by top-tier papers11
- ChatReID: Open-Ended Interactive Person Retrieval via Hierarchical Progressive Tuning for Vision Language ModelsKe Niu, Haiyang Yu, Mengyang Zhao, Teng Fu et al.ICCV 2025 · 5 citations
- Miss-ReID: Delivering Robust Multi-Modality Object Re-Identification Despite Missing ModalitiesRuida XiNeurIPS 2025 · 4 citations
- DisenQ: Disentangling Q-Former for Activity-BiometricsShehreen Azad, Yogesh Singh RawatICCV 2025 · 4 citations
- Prompt-Driven Transferable Adversarial Attack on Person Re-identification with Attribute-Aware Textual InversionYuan Bian, Min Liu, Yunqi Yi, Xueping Wang et al.ICCV 2025 · 3 citations
- Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-IdentificationKunlun Xu, Haotong Cheng, Jiangmeng Li, Xu Zou et al.CVPR 2026 · 2 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang et al.ICCV 2021 · 1,172 citations
- Omni-Scale Feature Learning for Person Re-IdentificationKaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao XiangICCV 2019 · 997 citations
- ABD-Net: Attentive but Diverse Person Re-IdentificationTianlong Chen, Shaojin Ding, Jingyi Xie, Ye Yuan et al.ICCV 2019 · 544 citations
Related papers
- Decoupled Identity and Attribute Tokenization for Person Re-IdentificationRui Shang, Min Liu, Xueping Wang, Yuan Bian et al.ACM MM 2025
- CLIP-driven View-aware Prompt Learning for Unsupervised Vehicle Re-identificationJiyang Xu, Qi Wang, Xin Xiong, Di Gai et al.AAAI 2025 · 8 citations
- CLIP-ReID: Exploiting Vision-Language Model for Image Re-identification without Concrete Text LabelsSiyuan Li, Li Sun, Qingli LiAAAI 2023 · 355 citations
- Unveiling the Power of CLIP in Unsupervised Visible-Infrared Person Re-IdentificationZhong Chen, Zhizhong Zhang, Xin Tan, Yanyun Qu et al.ACM MM 2023 · 65 citations
- ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-IdentificationCan Cui, Siteng Huang, Wenxuan Song, Pengxiang Ding et al.ACM MM 2024 · 18 citations
