Prototypical Prompting for Text-to-image Person Re-identification
Shuanglin Yan, Jun Liu, Neng Dong, Liyan Zhang, Jinhui Tang
摘要
In this paper, we study the problem of Text-to-Image Person Re-identification (TIReID), which aims to find images of the same identity described by a text sentence from a pool of candidate images. Benefiting from Vision-Language Pre-training, such as CLIP (Contrastive Language-Image Pretraining), the TIReID techniques have achieved remarkable progress recently. However, most existing methods only focus on instance-level matching and ignore identity-level matching, which involves associating multiple images and texts belonging to the same person. In this paper, we propose a novel prototypical prompting framework (Propot) designed to simultaneously model instance-level and identity-level matching for TIReID. Our Propot transforms the identity-level matching problem into a prototype learning problem, aiming to learn identity-enriched prototypes. Specifically, Propot works by 'initialize, adapt, enrich, then aggregate'. We first use CLIP to generate high-quality initial prototypes. Then, we propose a domain-conditional prototypical prompting (DPP) module to adapt the prototypes to the TIReID task using task-related information. Further, we propose an instance-conditional prototypical prompting (IPP) module to update prototypes conditioned on intra-modal and inter-modal instances to ensure prototype diversity. Finally, we design an adaptive prototype aggregation module to aggregate these prototypes, generating final identity-enriched prototypes. With identity-enriched prototypes, we diffuse its rich identity information to instances through prototype-to-instance contrastive loss to facilitate identity-level matching. Extensive experiments conducted on three benchmarks demonstrate the superiority of Propot compared to existing TIReID methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image MatchingYafei Zhang, Yongle Shang, Huafeng LiACM MM 2025 · 被引用 5 次
- Hierarchical Prompt Learning for Image- and Text-Based Person Re-IdentificationLinhan Zhou, Shuang Li, Neng Dong, Yonghang Tai 等AAAI 2026 · 被引用 4 次
- Text-based Aerial-Ground Person RetrievalXinyu Zhou, Yu Wu, Jiayao Ma, Wenhao Wang 等AAAI 2026 · 被引用 1 次
- Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person RetrievalTianlu Zheng, Yifan Zhang, Xiang An, Ziyong Feng 等EMNLP 2025 · 被引用 1 次
- Modeling Thousands of Human Annotators for Generalizable Text-to-Image Person Re-identificationJiayu Jiang, Changxing Ding, Wentao Tan, Junhong Wang 等CVPR 2025
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
相关 Paper
- Unveiling the Power of CLIP in Unsupervised Visible-Infrared Person Re-IdentificationZhong Chen, Zhizhong Zhang, Xin Tan, Yanyun Qu 等ACM MM 2023 · 被引用 65 次
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su 等AAAI 2025 · 被引用 23 次
- CLIP-driven View-aware Prompt Learning for Unsupervised Vehicle Re-identificationJiyang Xu, Qi Wang, Xin Xiong, Di Gai 等AAAI 2025 · 被引用 8 次
- Unified Pre-training with Pseudo Texts for Text-To-Image Person Re-identificationZhiyin Shao, Xinyu Zhang, Changxing Ding, Jian Wang 等ICCV 2023 · 被引用 46 次
- Prompt-Driven Referring Image Segmentation with Instance ContrastingChao Shang, Zichen Song, Heqian Qiu, Lanxiao Wang 等CVPR 2024 · 被引用 20 次
