ProtoHPE: Prototype-guided High-frequency Patch Enhancement for Visible-Infrared Person Re-identification
Guiwei Zhang, Yongfei Zhang, Zichang Tan
Abstract
Visible-Infrared person re-identification is challenging due to the large modality gap. To bridge the gap, most studies heavily rely on the correlation of visible-infrared holistic person images, which may perform poorly under severe distribution shifts. In contrast, we find that some cross-modal correlated high-frequency components contain discriminative visual patterns and are less affected by variations such as wavelength, pose, and background clutter than holistic images. Therefore, we are motivated to bridge the modality gap based on such high-frequency components, and propose Prototype-guided High-frequency Patch Enhancement (ProtoHPE) with two core designs. First, to enhance the representation ability of cross-modal correlated high-frequency components, we split patches with such components by Wavelet Transform and exponential moving average Vision Transformer (ViT), then empower ViT to take the split patches as auxiliary input. Second, to obtain semantically compact and discriminative high-frequency representations of the same identity, we propose Multimodal Prototypical Contrast. To be specific, it hierarchically captures comprehensive semantics of different modal instances, facilitating the aggregation of high-frequency representations belonging to the same identity. With it, ViT can capture key high-frequency components during inference without relying on ProtoHPE, thus bringing no extra complexity. Extensive experiments validate the effectiveness of ProtoHPE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f7d8ef5-493f-4956-a6bb-cc5efb23de74Cited by top-tier papers4
- Robust Pseudo-label Learning with Neighbor Relation for Unsupervised Visible-Infrared Person Re-IdentificationXiangbo Yin, Jiangming Shi, Yachao Zhang, Yang Lu et al.ACM MM 2024 · 28 citations
- BIT: Matching-based Bi-directional Interaction Transformation Network for Visible-Infrared Person Re-IdentificationHaoxuan Xu, Guanglin NiuCVPR 2026 · 3 citations
- MFEN: Multi-Frequency Expert Network for Visible-Infrared Person Re-IDXulin Li, Yan Lu, Bin Liu, Qinhong Yang et al.CVPR 2026 · 2 citations
- From Poses to Identity: Training-Free Person Re-Identification via Feature CentralizationChao Yuan, Guiwei Zhang, Changxiao Ma, Tianyi Zhang et al.CVPR 2025
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- PHA: Patch-Wise High-Frequency Augmentation for Transformer-Based Person Re-IdentificationGuiwei Zhang, Yongfei Zhang, Tianyu Zhang, Bo Li et al.CVPR 2023
- Modality Unifying Network for Visible-Infrared Person Re-IdentificationHao Yu, Xu Cheng, Wei Peng, Weihao Liu et al.ICCV 2023 · 75 citations
- Learning Progressive Modality-Shared Transformers for Effective Visible-Infrared Person Re-identificationHu Lu, Xuezhang Zou, Pingping ZhangAAAI 2023 · 183 citations
- X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-IdentificationChenyang Yu, Xuehu Liu, Pingping Zhang, Huchuan LuAAAI 2026 · 3 citations
- Learning by Aligning: Visible-Infrared Person Re-identification using Cross-Modal CorrespondencesHyunjong Park, Sanghoon Lee, Junghyup Lee, Bumsub HamICCV 2021 · 248 citations
