PLIP: Language-Image Pre-training for Person Representation Learning
Jialong Zuo, Jiahao Hong, Feng Zhang, Changqian Yu, Hanyu Zhou, Changxin Gao, Nong Sang, Jingdong Wang
Abstract
Language-image pre-training is an effective technique for learning powerful representations in general domains. However, when directly turning to person representation learning, these general pre-training methods suffer from unsatisfactory performance. The reason is that they neglect critical person-related characteristics, i.e., fine-grained attributes and identities. To address this issue, we propose a novel language-image pre-training framework for person representation learning, termed PLIP. Specifically, we elaborately design three pretext tasks: 1) Text-guided Image Colorization, aims to establish the correspondence between the person-related image regions and the fine-grained color-part textual phrases. 2) Image-guided Attributes Prediction, aims to mine fine-grained attribute information of the person body in the image; and 3) Identity-based Vision-Language Contrast, aims to correlate the cross-modal representations at the identity level rather than the instance level. Moreover, to implement our pre-train framework, we construct a large-scale person dataset with image-text pairs named SYNTH-PEDES by automatically generating textual annotations. We pre-train PLIP on SYNTH-PEDES and evaluate our models by spanning downstream person-centric tasks. PLIP not only significantly improves existing methods on all these tasks, but also shows great ability in the zero-shot and domain generalization settings. The code, dataset and weights will be released at https://github.com/Zplusdragon/PLIP
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42cbf259-809f-4d00-b5e6-9c1654bb1f14Cited by top-tier papers21
- UFineBench: Towards Text-based Person Retrieval with Ultra-fine GranularityJialong Zuo, Hanyu Zhou, Ying Nie, Feng Zhang et al.CVPR 2024 · 45 citations
- Robust Pseudo-label Learning with Neighbor Relation for Unsupervised Visible-Infrared Person Re-IdentificationXiangbo Yin, Jiangming Shi, Yachao Zhang, Yang Lu et al.ACM MM 2024 · 28 citations
- Pedestrian Attribute Recognition: A New Benchmark Dataset and a Large Language Model Augmented FrameworkJiandong Jin, Xiao Wang, Qian Zhu, Haiyang Wang et al.AAAI 2025 · 19 citations
- Cross-video Identity Correlating for Person Re-identification Pre-trainingJialong Zuo, Ying Nie, Hanyu Zhou, Huaxin Zhang et al.NeurIPS 2024 · 15 citations
- ReID5o: Achieving Omni Multi-modal Person Re-identification in a Single ModelJialong Zuo, Yongtai Deng, Mengdan Tan, Rui Jin et al.NeurIPS 2025 · 11 citations
Builds on42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Unified Pre-training with Pseudo Texts for Text-To-Image Person Re-identificationZhiyin Shao, Xinyu Zhang, Changxing Ding, Jian Wang et al.ICCV 2023 · 46 citations
- Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person RetrievalTianlu Zheng, Yifan Zhang, Xiang An, Ziyong Feng et al.EMNLP 2025 · 1 citation
- An Empirical Study of CLIP for Text-Based Person SearchMin Cao, Yang Bai, Ziyin Zeng, Mang Ye et al.AAAI 2024 · 111 citations
- Prototypical Prompting for Text-to-image Person Re-identificationShuanglin Yan, Jun Liu, Neng Dong, Liyan Zhang et al.ACM MM 2024 · 16 citations
- Towards Unified Text-based Person Retrieval: A Large-scale Multi-Attribute and Language Search BenchmarkShuyu Yang, Yinan Zhou, Zhedong Zheng, Yaxiong Wang et al.ACM MM 2023 · 162 citations
