LVLM-Driven Attribute-Aware Modeling for Visible-Infrared Person Re-Identification
Zhiqi Pang, Lingling Zhao, Junjie Wang, Chunyu Wang
摘要
Visible-infrared person re-identification (VI-ReID) aims to match visible and infrared images of the same individual. Supervised VI-ReID (SVI-ReID) methods have achieved promising performance under the guidance of manually annotated identity labels. However, the substantial annotation cost severely limits their scalability in real-world applications. As a result, unsupervised VI-ReID (UVI-ReID) methods have attracted increasing attention. These methods typically rely on pseudo-labels generated by clustering and matching algorithms to replace manual annotations. Nevertheless, the quality of pseudo-labels is often difficult to guarantee, and low-quality pseudo-labels can significantly hinder model performance improvements. To address these challenges, we explore the use of attribute arrays extracted by a large vision-language model (LVLM) to enhance VI-ReID, and propose a novel LVLM-driven attribute-aware modeling (LVLM-AAM) approach. Specifically, we first design an attribute-aware reliable labeling strategy, which refines intra-modality clustering results based on image-level attributes and improves inter-modality matching by grouping clusters according to cluster-level attributes. Next, we develop an explicit-implicit attribute fusion module, which integrates explicit and implicit attributes to obtain more fine-grained identity-related text features. Finally, we introduce an attribute-aware contrastive learning module, which jointly leverages static and dynamic text features to promote modality-invariant feature learning. Extensive experiments conducted on VI-ReID datasets validate the effectiveness of the proposed LVLM-AAM and its individual components. LVLM-AAM not only significantly outperforms existing unsupervised methods but also surpasses several supervised methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li 等AAAI 2020 · 被引用 4,134 次
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang 等ICCV 2021 · 被引用 1,172 次
- CLIP-ReID: Exploiting Vision-Language Model for Image Re-identification without Concrete Text LabelsSiyuan Li, Li Sun, Qingli LiAAAI 2023 · 被引用 355 次
相关 Paper
- Empowering Visible-Infrared Person Re-Identification with Large Foundation ModelsZhangyi Hu, Bin Yang, Mang YeNeurIPS 2024 · 被引用 45 次
- Augmented Dual-Contrastive Aggregation Learning for Unsupervised Visible-Infrared Person Re-IdentificationBin Yang, Mang Ye, Jun Chen, Zesen WuACM MM 2022 · 被引用 113 次
- Unveiling the Power of CLIP in Unsupervised Visible-Infrared Person Re-IdentificationZhong Chen, Zhizhong Zhang, Xin Tan, Yanyun Qu 等ACM MM 2023 · 被引用 65 次
- Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-IdentificationKunlun Xu, Haotong Cheng, Jiangmeng Li, Xu Zou 等CVPR 2026 · 被引用 2 次
- Efficient Bilateral Cross-Modality Cluster Matching for Unsupervised Visible-Infrared Person ReIDDe Cheng, Lingfeng He, Nannan Wang, Shizhou Zhang 等ACM MM 2023 · 被引用 36 次
