Unveiling the Power of CLIP in Unsupervised Visible-Infrared Person Re-Identification
Zhong Chen, Zhizhong Zhang, Xin Tan, Yanyun Qu, Yuan Xie
Abstract
Large-scale Vision-Language Pre-training (VLP) model, e.g., CLIP, has demonstrated its natural advantage in generating textual descriptions for images. These textual descriptions afford us greater semantic monitoring insights while not requiring any domain knowledge. In this paper, we propose a new prompt learning paradigm for unsupervised visible-infrared person re-identification (USL-VI-ReID) by taking full advantage of the visual-text representation ability from CLIP. In our framework, we establish a learnable cluster-aware prompt for person images and obtain textual descriptions allowing for subsequent unsupervised training. This description complements the rigid pseudo-labels and provides an important semantic supervised signal. On that basis, we propose a new memory-swapping contrastive learning, where we first find the correlated cross-modal prototypes by the Hungarian matching method and then swap the prototype pairs in the memory. Thus typical contrastive learning without any change could easily associate the cross-modal information. Extensive experiments on the benchmark datasets demonstrate the effectiveness of our method. For example, on SYSU-MM01 we arrive at 54.0% in terms of Rank-1 accuracy, over 9% improvement against state-of-the-art approaches. Code is available at https://github.com/CzAngus/CCLNet.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers16
- Learning Commonality, Divergence and Variety for Unsupervised Visible-Infrared Person Re-identificationJiangming Shi, Xiangbo Yin, Yachao Zhang, Zhizhong Zhang et al.NeurIPS 2024 · 36 citations
- Robust Pseudo-label Learning with Neighbor Relation for Unsupervised Visible-Infrared Person Re-IdentificationXiangbo Yin, Jiangming Shi, Yachao Zhang, Yang Lu et al.ACM MM 2024 · 28 citations
- Weakly Supervised Visible-Infrared Person Re-Identification via Heterogeneous Expert Collaborative Consistency LearningYafei Zhang, Lingqi Kong, Huafeng Li, Jie WenICCV 2025 · 10 citations
- Relieving Universal Label Noise for Unsupervised Visible-Infrared Person Re-Identification by Inferring from NeighborsXiao Teng, Long Lan, Dingyao Chen, Kele Xu et al.AAAI 2025 · 4 citations
- Learning Source-Free Domain Adaptation for Visible-Infrared Person Re-IdentificationYongxiang Li, Yanglin Feng, Yuan Sun, Dezhong Peng et al.NeurIPS 2025 · 4 citations
Related papers
- Prototypical Prompting for Text-to-image Person Re-identificationShuanglin Yan, Jun Liu, Neng Dong, Liyan Zhang et al.ACM MM 2024 · 16 citations
- Unified Pre-training with Pseudo Texts for Text-To-Image Person Re-identificationZhiyin Shao, Xinyu Zhang, Changxing Ding, Jian Wang et al.ICCV 2023 · 46 citations
- CLIP-driven View-aware Prompt Learning for Unsupervised Vehicle Re-identificationJiyang Xu, Qi Wang, Xin Xiong, Di Gai et al.AAAI 2025 · 8 citations
- An Empirical Study of CLIP for Text-Based Person SearchMin Cao, Yang Bai, Ziyin Zeng, Mang Ye et al.AAAI 2024 · 111 citations
- CLIP-ReID: Exploiting Vision-Language Model for Image Re-identification without Concrete Text LabelsSiyuan Li, Li Sun, Qingli LiAAAI 2023 · 355 citations
