ProtoHPE: Prototype-guided High-frequency Patch Enhancement for Visible-Infrared Person Re-identification
Guiwei Zhang, Yongfei Zhang, Zichang Tan
摘要
Visible-Infrared person re-identification is challenging due to the large modality gap. To bridge the gap, most studies heavily rely on the correlation of visible-infrared holistic person images, which may perform poorly under severe distribution shifts. In contrast, we find that some cross-modal correlated high-frequency components contain discriminative visual patterns and are less affected by variations such as wavelength, pose, and background clutter than holistic images. Therefore, we are motivated to bridge the modality gap based on such high-frequency components, and propose Prototype-guided High-frequency Patch Enhancement (ProtoHPE) with two core designs. First, to enhance the representation ability of cross-modal correlated high-frequency components, we split patches with such components by Wavelet Transform and exponential moving average Vision Transformer (ViT), then empower ViT to take the split patches as auxiliary input. Second, to obtain semantically compact and discriminative high-frequency representations of the same identity, we propose Multimodal Prototypical Contrast. To be specific, it hierarchically captures comprehensive semantics of different modal instances, facilitating the aggregation of high-frequency representations belonging to the same identity. With it, ViT can capture key high-frequency components during inference without relying on ProtoHPE, thus bringing no extra complexity. Extensive experiments validate the effectiveness of ProtoHPE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Robust Pseudo-label Learning with Neighbor Relation for Unsupervised Visible-Infrared Person Re-IdentificationXiangbo Yin, Jiangming Shi, Yachao Zhang, Yang Lu 等ACM MM 2024 · 被引用 28 次
- BIT: Matching-based Bi-directional Interaction Transformation Network for Visible-Infrared Person Re-IdentificationHaoxuan Xu, Guanglin NiuCVPR 2026 · 被引用 3 次
- MFEN: Multi-Frequency Expert Network for Visible-Infrared Person Re-IDXulin Li, Yan Lu, Bin Liu, Qinhong Yang 等CVPR 2026 · 被引用 2 次
- From Poses to Identity: Training-Free Person Re-Identification via Feature CentralizationChao Yuan, Guiwei Zhang, Changxiao Ma, Tianyi Zhang 等CVPR 2025
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
相关 Paper
- PHA: Patch-Wise High-Frequency Augmentation for Transformer-Based Person Re-IdentificationGuiwei Zhang, Yongfei Zhang, Tianyu Zhang, Bo Li 等CVPR 2023
- Modality Unifying Network for Visible-Infrared Person Re-IdentificationHao Yu, Xu Cheng, Wei Peng, Weihao Liu 等ICCV 2023 · 被引用 75 次
- Learning Progressive Modality-Shared Transformers for Effective Visible-Infrared Person Re-identificationHu Lu, Xuezhang Zou, Pingping ZhangAAAI 2023 · 被引用 183 次
- X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-IdentificationChenyang Yu, Xuehu Liu, Pingping Zhang, Huchuan LuAAAI 2026 · 被引用 3 次
- Learning by Aligning: Visible-Infrared Person Re-identification using Cross-Modal CorrespondencesHyunjong Park, Sanghoon Lee, Junghyup Lee, Bumsub HamICCV 2021 · 被引用 248 次
