Image-Text Knowledge Modeling for Unsupervised Multi-Scenario Person Re-Identification
Zhiqi Pang, Lingling Zhao, Yang Liu, Chunyu Wang, Gaurav Sharma
Abstract
We propose unsupervised multi-scenario (UMS) person reidentification (ReID) as a new task that expands ReID across diverse scenarios (cross-resolution, clothing change, etc.) within a single coherent framework. To tackle UMS-ReID, we introduce image-text knowledge modeling (ITKM) -a three-stage framework that effectively exploits the representational power of vision-language models. We start with a pre-trained CLIP model with an image encoder and a text encoder. In Stage I, we introduce a scenario embedding in the image encoder and fine-tune the encoder to adaptively leverage knowledge from multiple scenarios. In Stage II, we optimize a set of learned text embeddings to associate with pseudo-labels from Stage I and introduce a multi-scenario separation loss to increase the divergence between interscenario text representations. In Stage III, we first introduce cluster-level and instance-level heterogeneous matching modules to obtain reliable heterogeneous positive pairs (e.g., a visible image and an infrared image of the same person) within each scenario. Next, we propose a dynamic text representation update strategy to maintain consistency between text and image supervision signals. Experimental results across multiple scenarios demonstrate the superiority and generalizability of ITKM; it not only outperforms existing scenario-specific methods but also enhances overall performance by integrating knowledge from multiple scenarios. * This is a preprint of a paper accepted for AAAI 2026 that is expanded to include Supplementary Material. Copyright will transfer to AAAI for the published paper (Pang et al. 2026 ), which should be cited for referencing the work presented here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b642bca-a9d4-4f8a-bdc7-39774f7e74efBuilds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Part-based Pseudo Label Refinement for Unsupervised Person Re-identificationYoonki Cho, Woo Jae Kim, Seunghoon Hong, Sung-Eui YoonCVPR 2022 · 271 citations
- ICE: Inter-instance Contrastive Encoding for Unsupervised Person Re-identificationHao Chen, Benoit Lagadec, François BrémondICCV 2021 · 258 citations
- Camera-Aware Proxies for Unsupervised Person Re-IdentificationMenglin Wang, Baisheng Lai, Jianqiang Huang, Xiaojin Gong et al.AAAI 2021 · 247 citations
Related papers
- Unveiling the Power of CLIP in Unsupervised Visible-Infrared Person Re-IdentificationZhong Chen, Zhizhong Zhang, Xin Tan, Yanyun Qu et al.ACM MM 2023 · 65 citations
- Identity-Clothing Similarity Modeling for Unsupervised Clothing Change Person Re-IdentificationZhiqi Pang, Junjie Wang, Lingling Zhao, Chunyu WangCVPR 2025
- CLIP-driven View-aware Prompt Learning for Unsupervised Vehicle Re-identificationJiyang Xu, Qi Wang, Xin Xiong, Di Gai et al.AAAI 2025 · 8 citations
- Prototypical Prompting for Text-to-image Person Re-identificationShuanglin Yan, Jun Liu, Neng Dong, Liyan Zhang et al.ACM MM 2024 · 16 citations
- CLIMB-ReID: A Hybrid CLIP-Mamba Framework for Person Re-IdentificationChenyang Yu, Xuehu Liu, Jiawen Zhu, Yuhao Wang et al.AAAI 2025 · 17 citations
