VRCLIP: Multimodal Canonical Correlation Alignment for CLIP-Driven Vision-Radio Person Re-Identification
Rui Zhang, Yaqi Wang, Yadong Li, Ruixu Geng, Jianyang Wang, Qijun Ying, Dongheng Zhang, Yang Hu, Yan Chen
Abstract
Person re-identification (ReID) is critical for public safety, yet the performance of RGB-based methods is limited under challenging lighting and occlusion conditions. In contrast, low-frequency radio frequency (RF) signals, with their superior penetration capability and illumination invariance, provide ideal complementary information. However, a key challenge in fusing these heterogeneous modalities lies in the conventional approach that relies heavily on cross‑modal distribution matching, which often over‑regularizes and weakens the discriminative capacity within each modality. Rather than enforcing direct distribution alignment, canonical correlation analysis (CCA) constructs a shared subspace that maximizes cross‑modal correlation, inherently balancing modality specificity and shared semantics. Inspired by this, we reformulate cross-modal alignment as a correlation maximization problem, avoiding direct constraints on feature distributions and guiding the model to harmonize intra‑modal discriminative learning with cross‑modal alignment. Specifically, VRCLIP first refines CLIP’s visual encoder with illumination‑disentangling objectives, then aligns RGB and RF embeddings in a canonical correlation subspace, and finally employs an RF‑anchored reliability gate for adaptive fusion. To advance the area, we will release VRR, the first large‑scale vision–radio ReID dataset with over 650K paired image–radar samples and position annotations for 31 participants. Extensive experiments show state‑of‑the‑art 93.9% mAP and robust generalization across diverse lighting and occlusion conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- PointMamba: A Simple State Space Model for Point Cloud AnalysisDingkang Liang, Xin Zhou, Wei Xu, Xingkui Zhu et al.NeurIPS 2024 · 380 citations
- Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identificationYongming Rao, Guangyi Chen, Jiwen Lu, Jie ZhouICCV 2021 · 330 citations
- Clothes-Changing Person Re-identification with RGB Modality OnlyXinqian Gu, Hong Chang, Bingpeng Ma, Shutao Bai et al.CVPR 2022 · 226 citations
Related papers
- X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-IdentificationChenyang Yu, Xuehu Liu, Pingping Zhang, Huchuan LuAAAI 2026 · 3 citations
- Spatial-Frequency Collaborative Learning for Occluded Visible-Infrared Person Re-IdentificationJIan Yu, Yujian Feng, Shuai You, Zhongkai Zhou et al.CVPR 2026
- Learning Concordant Attention via Target-aware Alignment for Visible-Infrared Person Re-identificationJianbing Wu, Hong Liu, Yuxin Su, Wei Shi et al.ICCV 2023 · 45 citations
- Learning Modal-Invariant and Temporal-Memory for Video-based Visible-Infrared Person Re-IdentificationXinyu Lin, Jinxing Li, Zeyu Ma, Huafeng Li et al.CVPR 2022 · 81 citations
- Cross Vision-RF Gait Re-identification with Low-cost RGB-D Cameras and mmWave RadarsDongjiang Cao, Ruofeng Liu, Hao Li, Shuai Wang et al.UbiComp 2022 · 47 citations
