X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-Identification
Chenyang Yu, Xuehu Liu, Pingping Zhang, Huchuan Lu
Abstract
Large-scale vision-language models (e.g., CLIP) have recently achieved remarkable performance in retrieval tasks, yet their potential for Video-based Visible-Infrared Person Re-Identification (VVI-ReID) remains largely unexplored. The primary challenges are narrowing the modality gap and leveraging spatiotemporal information in video sequences. To address the above issues, in this paper, we propose a novel crossmodality feature learning framework named X-ReID for VVI-ReID. Specifically, we first propose a Cross-modality Prototype Collaboration (CPC) to align and integrate features from different modalities, guiding the network to reduce the modality discrepancy. Then, a Multi-granularity Information Interaction (MII) is designed, incorporating short-term interactions from adjacent frames, long-term cross-frame information fusion, and cross-modality feature alignment to enhance temporal modeling and further reduce modality gaps. Finally, by integrating multi-granularity information, a robust sequence-level representation is achieved. Extensive experiments on two large-scale VVI-ReID benchmarks (i.e., HITSZ-VCM and BUPTCampus) demonstrate the superiority of our method over state-of-the-art methods. The source code is released at https://github.com/AsuradaYuci/X-ReID .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1de431f9-72bf-4640-a6bb-d8fa29ca31b9Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- RGB-Infrared Cross-Modality Person Re-Identification via Joint Pixel and Feature AlignmentGuan'an Wang, Tianzhu Zhang, Jian Cheng, Si Liu et al.ICCV 2019 · 464 citations
- CLIP-ReID: Exploiting Vision-Language Model for Image Re-identification without Concrete Text LabelsSiyuan Li, Li Sun, Qingli LiAAAI 2023 · 355 citations
Related papers
- Learning by Aligning: Visible-Infrared Person Re-identification using Cross-Modal CorrespondencesHyunjong Park, Sanghoon Lee, Junghyup Lee, Bumsub HamICCV 2021 · 248 citations
- Video-based Visible-Infrared Person Re-Identification via Style Disturbance Defense and Dual InteractionChuhao Zhou, Jinxing Li, Huafeng Li, Guangming Lu et al.ACM MM 2023 · 23 citations
- CLIMB-ReID: A Hybrid CLIP-Mamba Framework for Person Re-IdentificationChenyang Yu, Xuehu Liu, Jiawen Zhu, Yuhao Wang et al.AAAI 2025 · 17 citations
- Cross-Modality Person Re-identification with Memory-Based Contrastive EmbeddingDe Cheng, Xiaolong Wang, Nannan Wang, Zhen Wang et al.AAAI 2023 · 22 citations
- Hyperbolic Hierarchical Alignment for Video-Based Visible-Infrared Person Re-IdentificationShuang Li, Changjiang Kuang, Jiaxu Leng, Mingpi Tan et al.ICML 2026
