Prototypical Prompting for Text-to-image Person Re-identification
Shuanglin Yan, Jun Liu, Neng Dong, Liyan Zhang, Jinhui Tang
Abstract
In this paper, we study the problem of Text-to-Image Person Re-identification (TIReID), which aims to find images of the same identity described by a text sentence from a pool of candidate images. Benefiting from Vision-Language Pre-training, such as CLIP (Contrastive Language-Image Pretraining), the TIReID techniques have achieved remarkable progress recently. However, most existing methods only focus on instance-level matching and ignore identity-level matching, which involves associating multiple images and texts belonging to the same person. In this paper, we propose a novel prototypical prompting framework (Propot) designed to simultaneously model instance-level and identity-level matching for TIReID. Our Propot transforms the identity-level matching problem into a prototype learning problem, aiming to learn identity-enriched prototypes. Specifically, Propot works by 'initialize, adapt, enrich, then aggregate'. We first use CLIP to generate high-quality initial prototypes. Then, we propose a domain-conditional prototypical prompting (DPP) module to adapt the prototypes to the TIReID task using task-related information. Further, we propose an instance-conditional prototypical prompting (IPP) module to update prototypes conditioned on intra-modal and inter-modal instances to ensure prototype diversity. Finally, we design an adaptive prototype aggregation module to aggregate these prototypes, generating final identity-enriched prototypes. With identity-enriched prototypes, we diffuse its rich identity information to instances through prototype-to-instance contrastive loss to facilitate identity-level matching. Extensive experiments conducted on three benchmarks demonstrate the superiority of Propot compared to existing TIReID methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc25da80-10e7-4d15-b712-bb6371242885Cited by top-tier papers6
- Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image MatchingYafei Zhang, Yongle Shang, Huafeng LiACM MM 2025 · 5 citations
- Hierarchical Prompt Learning for Image- and Text-Based Person Re-IdentificationLinhan Zhou, Shuang Li, Neng Dong, Yonghang Tai et al.AAAI 2026 · 4 citations
- Text-based Aerial-Ground Person RetrievalXinyu Zhou, Yu Wu, Jiayao Ma, Wenhao Wang et al.AAAI 2026 · 1 citation
- Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person RetrievalTianlu Zheng, Yifan Zhang, Xiang An, Ziyong Feng et al.EMNLP 2025 · 1 citation
- Modeling Thousands of Human Annotators for Generalizable Text-to-Image Person Re-identificationJiayu Jiang, Changxing Ding, Wentao Tan, Junhong Wang et al.CVPR 2025
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
Related papers
- Unveiling the Power of CLIP in Unsupervised Visible-Infrared Person Re-IdentificationZhong Chen, Zhizhong Zhang, Xin Tan, Yanyun Qu et al.ACM MM 2023 · 65 citations
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su et al.AAAI 2025 · 23 citations
- CLIP-driven View-aware Prompt Learning for Unsupervised Vehicle Re-identificationJiyang Xu, Qi Wang, Xin Xiong, Di Gai et al.AAAI 2025 · 8 citations
- Unified Pre-training with Pseudo Texts for Text-To-Image Person Re-identificationZhiyin Shao, Xinyu Zhang, Changxing Ding, Jian Wang et al.ICCV 2023 · 46 citations
- Prompt-Driven Referring Image Segmentation with Instance ContrastingChao Shang, Zichen Song, Heqian Qiu, Lanxiao Wang et al.CVPR 2024 · 20 citations
