TF-CLIP: Learning Text-Free CLIP for Video-Based Person Re-identification
Chenyang Yu, Xuehu Liu, Yingquan Wang, Pingping Zhang, Huchuan Lu
Abstract
Large-scale language-image pre-trained models (e.g., CLIP) have shown superior performances on many cross-modal retrieval tasks. However, the problem of transferring the knowledge learned from such models to video-based person re-identification (ReID) has barely been explored. In addition, there is a lack of decent text descriptions in current ReID benchmarks. To address these issues, in this work, we propose a novel one-stage text-free CLIP-based learning framework named TF-CLIP for video-based person ReID. More specifically, we extract the identity-specific sequence feature as the CLIP-Memory to replace the text feature. Meanwhile, we design a Sequence-Specific Prompt (SSP) module to update the CLIP-Memory online. To capture temporal information, we further propose a Temporal Memory Diffusion (TMD) module, which consists of two key components: Temporal Memory Construction (TMC) and Memory Diffusion (MD). Technically, TMC allows the frame-level memories in a sequence to communicate with each other, and to extract temporal information based on the relations within the sequence. MD further diffuses the temporal memories to each token in the original features to obtain more robust sequence features. Extensive experiments demonstrate that our proposed method shows much better results than other state-of-the-art methods on MARS, LS-VID and iLIDS-VID.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d71d51ec-d528-4743-8af3-1f8573b236b9Cited by top-tier papers23
- Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-IdentificationPingping Zhang, Yuhao Wang, Yang Liu, Zhengzheng Tu et al.CVPR 2024 · 42 citations
- Learning Commonality, Divergence and Variety for Unsupervised Visible-Infrared Person Re-identificationJiangming Shi, Xiangbo Yin, Yachao Zhang, Zhizhong Zhang et al.NeurIPS 2024 · 36 citations
- DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-IdentificationYuhao Wang, Yang Liu, Aihua Zheng, Pingping ZhangAAAI 2025 · 31 citations
- MambaPro: Multi-Modal Object Re-identification with Mamba Aggregation and Synergistic PromptYuhao Wang, Xuehu Liu, Tianyu Yan, Yang Liu et al.AAAI 2025 · 30 citations
- Historical Test-time Prompt Tuning for Vision Foundation ModelsJingyi Zhang, Jiaxing Huang, Xiaoqin Zhang, Ling Shao et al.NeurIPS 2024 · 29 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
Related papers
- CLIMB-ReID: A Hybrid CLIP-Mamba Framework for Person Re-IdentificationChenyang Yu, Xuehu Liu, Jiawen Zhu, Yuhao Wang et al.AAAI 2025 · 17 citations
- Unveiling the Power of CLIP in Unsupervised Visible-Infrared Person Re-IdentificationZhong Chen, Zhizhong Zhang, Xin Tan, Yanyun Qu et al.ACM MM 2023 · 65 citations
- Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-IdentificationRifen Lin, Alex Jinpeng Wang, Jiawei Mo, Min LiAAAI 2026
- X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-IdentificationChenyang Yu, Xuehu Liu, Pingping Zhang, Huchuan LuAAAI 2026 · 3 citations
- CLIP-ReID: Exploiting Vision-Language Model for Image Re-identification without Concrete Text LabelsSiyuan Li, Li Sun, Qingli LiAAAI 2023 · 355 citations
