PHA: Patch-Wise High-Frequency Augmentation for Transformer-Based Person Re-Identification
Guiwei Zhang, Yongfei Zhang, Tianyu Zhang, Bo Li, Shiliang Pu
摘要
Although recent studies empirically show that injecting Convolutional Neural Networks (CNNs) into Vision Transformers (ViTs) can improve the performance of person reidentification, the rationale behind it remains elusive. From a frequency perspective, we reveal that ViTs perform worse than CNNs in preserving key high-frequency components (e.g, clothes texture details) since high-frequency components are inevitably diluted by low-frequency ones due to the intrinsic Self-Attention within ViTs. To remedy such inadequacy of the ViT, we propose a Patch-wise Highfrequency Augmentation (PHA) method with two core designs. First, to enhance the feature representation ability of high-frequency components, we split patches with highfrequency components by the Discrete Haar Wavelet Transform, then empower the ViT to take the split patches as auxiliary input. Second, to prevent high-frequency components from being diluted by low-frequency ones when taking the entire sequence as input during network optimization, we propose a novel patch-wise contrastive loss. From the view of gradient optimization, it acts as an implicit augmentation to improve the representation ability of key high-frequency components. This benefits the ViT to capture key highfrequency components to extract discriminative person representations. PHA is necessary during training and can be removed during inference, without bringing extra complexity. Extensive experiments on widely-used ReID datasets validate the effectiveness of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-IdentificationPingping Zhang, Yuhao Wang, Yang Liu, Zhengzheng Tu 等CVPR 2024 · 被引用 42 次
- A Pedestrian is Worth One Prompt: Towards Language Guidance Person Re- IdentificationZexian Yang, Dayan Wu, Chenming Wu, Zheng Lin 等CVPR 2024 · 被引用 26 次
- ProtoHPE: Prototype-guided High-frequency Patch Enhancement for Visible-Infrared Person Re-identificationGuiwei Zhang, Yongfei Zhang, Zichang TanACM MM 2023 · 被引用 23 次
- Heterogeneous Test-Time Training for Multi-Modal Person Re-identificationZi Wang, Huaibo Huang, Aihua Zheng, Ran HeAAAI 2024 · 被引用 22 次
- SEAS: ShapE-Aligned Supervision for Person Re-IdentificationHaidong Zhu, Pranav Budhwant, Zhaoheng Zheng, Ram NevatiaCVPR 2024 · 被引用 14 次
它引用的顶会 Paper21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang 等ICCV 2021 · 被引用 1,172 次
- Global Filter Networks for Image ClassificationYongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu 等NeurIPS 2021 · 被引用 798 次
- Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-IdentificationHaowei Zhu, Wenjing Ke, Dong Li, Ji Liu 等CVPR 2022 · 被引用 251 次
- Pose-Guided Feature Disentangling for Occluded Person Re-identification Based on TransformerTao Wang, Hong Liu, Pinhao Song, Tianyu Guo 等AAAI 2022 · 被引用 248 次
相关 Paper
- Understanding and Improving Robustness of Vision Transformers through Patch-based Negative AugmentationYao Qin, Chiyuan Zhang, Ting Chen, Balaji Lakshminarayanan 等NeurIPS 2022 · 被引用 68 次
- TransFace: Calibrating Transformer Training for Face Recognition from a Data-Centric PerspectiveJun Dan, Yang Liu, Haoyu Xie, Jiankang Deng 等ICCV 2023 · 被引用 36 次
- The Principle of Diversity: Training Stronger Vision Transformers Calls for Reducing All Levels of RedundancyTianlong Chen, Zhenyu Zhang, Yu Cheng, Ahmed Awadallah 等CVPR 2022 · 被引用 36 次
- Revisiting Vision Transformer from the View of Path EnsembleShuning Chang, Pichao Wang, Hao Luo, Fan Wang 等ICCV 2023 · 被引用 8 次
- Scalable Vision Transformers with Hierarchical PoolingZizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He 等ICCV 2021 · 被引用 154 次
