Empowering Visible-Infrared Person Re-Identification with Large Foundation Models
Zhangyi Hu, Bin Yang, Mang Ye
Abstract
Visible-Infrared Person Re-identification (VI-ReID) is a challenging cross-modal retrieval task due to significant modality differences, primarily resulting from the absence of color information in the infrared modality. The development of large foundation models like Large Language Models (LLMs) and Vision Language Models (VLMs) motivates us to explore a feasible solution to empower VI-ReID with off-the-shelf large foundation models. To this end, we propose a novel Text-enhanced VI-ReID framework driven by Large Foundation Models (TVI-LFM). The core idea is to enrich the representation of the infrared modality with textual descriptions automatically generated by VLMs. Specifically, we incorporate a pre-trained VLM to extract textual features from texts generated by VLM and augmented by LLM, and incrementally fine-tune the text encoder to minimize the domain gap between generated texts and original visual modalities. Meanwhile, to enhance the infrared modality with extracted textual representations, we leverage modality alignment capabilities of VLMs and VLM-generated feature-level filters. This enables the text model to learn complementary features from the infrared modality, ensuring the semantic structural consistency between the fusion modality and the visible modality. Furthermore, we introduce modality joint learning to align features across all modalities, ensuring that textual features maintain stable semantic representation of overall pedestrian appearance during complementary information learning. Additionally, a modality ensemble retrieval strategy is proposed to leverage complementary strengths of each query modality to improve retrieval effectiveness and robustness. Extensive experiments on three expanded VI-ReID datasets demonstrate that our method significantly improves the retrieval performance, paving the way for the utilization of large foundation models in downstream multi-modal retrieval tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f22dd38e-c742-45ba-aad9-9857381715ffCited by top-tier papers21
- DisenQ: Disentangling Q-Former for Activity-BiometricsShehreen Azad, Yogesh Singh RawatICCV 2025 · 4 citations
- BMW: Bidirectionally Memory bank reWriting for Unsupervised Person Re-IdentificationXiaobin Liu, Jianing Li, Baiwei Guo, Wenbin Zhu et al.NeurIPS 2025 · 3 citations
- Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-IdentificationKunlun Xu, Haotong Cheng, Jiangmeng Li, Xu Zou et al.CVPR 2026 · 2 citations
- Unbiased Prototype Consistency Learning for Multi-Modal and Multi-Task Object Re-IdentificationZhongao Zhou, Bin Yang, Wenke Huang, Jun Chen et al.NeurIPS 2025 · 2 citations
- Joint Implicit and Explicit Language Learning for Pedestrian Attribute RecognitionYukang Zhang, Lei Tan, Yang Lu, Yan Yan et al.AAAI 2026 · 1 citation
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- Miss-ReID: Delivering Robust Multi-Modality Object Re-Identification Despite Missing ModalitiesRuida XiNeurIPS 2025 · 4 citations
- LVLM-Driven Attribute-Aware Modeling for Visible-Infrared Person Re-IdentificationZhiqi Pang, Lingling Zhao, Junjie Wang, Chunyu WangNeurIPS 2025 · 1 citation
- Cross-Modal Semantic Decoupling and Transfer for Text-to-Visible-Infrared Person Re-IdentificationZiang Zhang, Bin Yang, Mang YeICML 2026
- X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-IdentificationChenyang Yu, Xuehu Liu, Pingping Zhang, Huchuan LuAAAI 2026 · 3 citations
- Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReIDWentao Tan, Changxing Ding, Jiayu Jiang, Fei Wang et al.CVPR 2024 · 34 citations
