PHA: Patch-Wise High-Frequency Augmentation for Transformer-Based Person Re-Identification
Guiwei Zhang, Yongfei Zhang, Tianyu Zhang, Bo Li, Shiliang Pu
Abstract
Although recent studies empirically show that injecting Convolutional Neural Networks (CNNs) into Vision Transformers (ViTs) can improve the performance of person reidentification, the rationale behind it remains elusive. From a frequency perspective, we reveal that ViTs perform worse than CNNs in preserving key high-frequency components (e.g, clothes texture details) since high-frequency components are inevitably diluted by low-frequency ones due to the intrinsic Self-Attention within ViTs. To remedy such inadequacy of the ViT, we propose a Patch-wise Highfrequency Augmentation (PHA) method with two core designs. First, to enhance the feature representation ability of high-frequency components, we split patches with highfrequency components by the Discrete Haar Wavelet Transform, then empower the ViT to take the split patches as auxiliary input. Second, to prevent high-frequency components from being diluted by low-frequency ones when taking the entire sequence as input during network optimization, we propose a novel patch-wise contrastive loss. From the view of gradient optimization, it acts as an implicit augmentation to improve the representation ability of key high-frequency components. This benefits the ViT to capture key highfrequency components to extract discriminative person representations. PHA is necessary during training and can be removed during inference, without bringing extra complexity. Extensive experiments on widely-used ReID datasets validate the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bfa93d0a-fa7c-405e-a676-bc0882abffd3Cited by top-tier papers19
- Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-IdentificationPingping Zhang, Yuhao Wang, Yang Liu, Zhengzheng Tu et al.CVPR 2024 · 42 citations
- A Pedestrian is Worth One Prompt: Towards Language Guidance Person Re- IdentificationZexian Yang, Dayan Wu, Chenming Wu, Zheng Lin et al.CVPR 2024 · 26 citations
- ProtoHPE: Prototype-guided High-frequency Patch Enhancement for Visible-Infrared Person Re-identificationGuiwei Zhang, Yongfei Zhang, Zichang TanACM MM 2023 · 23 citations
- Heterogeneous Test-Time Training for Multi-Modal Person Re-identificationZi Wang, Huaibo Huang, Aihua Zheng, Ran HeAAAI 2024 · 22 citations
- SEAS: ShapE-Aligned Supervision for Person Re-IdentificationHaidong Zhu, Pranav Budhwant, Zhaoheng Zheng, Ram NevatiaCVPR 2024 · 14 citations
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang et al.ICCV 2021 · 1,172 citations
- Global Filter Networks for Image ClassificationYongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu et al.NeurIPS 2021 · 798 citations
- Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-IdentificationHaowei Zhu, Wenjing Ke, Dong Li, Ji Liu et al.CVPR 2022 · 251 citations
- Pose-Guided Feature Disentangling for Occluded Person Re-identification Based on TransformerTao Wang, Hong Liu, Pinhao Song, Tianyu Guo et al.AAAI 2022 · 248 citations
Related papers
- Understanding and Improving Robustness of Vision Transformers through Patch-based Negative AugmentationYao Qin, Chiyuan Zhang, Ting Chen, Balaji Lakshminarayanan et al.NeurIPS 2022 · 68 citations
- TransFace: Calibrating Transformer Training for Face Recognition from a Data-Centric PerspectiveJun Dan, Yang Liu, Haoyu Xie, Jiankang Deng et al.ICCV 2023 · 36 citations
- The Principle of Diversity: Training Stronger Vision Transformers Calls for Reducing All Levels of RedundancyTianlong Chen, Zhenyu Zhang, Yu Cheng, Ahmed Awadallah et al.CVPR 2022 · 36 citations
- Revisiting Vision Transformer from the View of Path EnsembleShuning Chang, Pichao Wang, Hao Luo, Fan Wang et al.ICCV 2023 · 8 citations
- Scalable Vision Transformers with Hierarchical PoolingZizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He et al.ICCV 2021 · 154 citations
