HAT: Hierarchical Aggregation Transformers for Person Re-identification
Guowen Zhang, Pingping Zhang, Jinqing Qi, Huchuan Lu
Abstract
Recently, with the advance of deep Convolutional Neural Networks (CNNs), person Re-Identification (Re-ID) has witnessed great success in various applications. However, with limited receptive fields of CNNs, it is still challenging to extract discriminative representations in a global view for persons under non-overlapped cameras. Meanwhile, Transformers demonstrate strong abilities of modeling long-range dependencies for spatial and sequential data. In this work, we take advantages of both CNNs and Transformers, and propose a novel learning framework named Hierarchical Aggregation Transformer (HAT) for image-based person Re-ID with high performance. To achieve this goal, we first propose a Deeply Supervised Aggregation (DSA) to recurrently aggregate hierarchical features from CNN backbones. With multi-granularity supervisions, the DSA can enhance multi-scale features for person retrieval, which is very different from previous methods. Then, we introduce a Transformer-based Feature Calibration (TFC) to integrate low-level detail information as the global prior for high-level semantic information. The proposed TFC is inserted to each level of hierarchical features, resulting in great performance improvements. To our best knowledge, this work is the first to take advantages of both CNNs and Transformers for image-based person Re-ID. Comprehensive experiments on four large-scale Re-ID benchmarks demonstrate that our method shows better results than several state-of-the-art methods. The code is released at https://github.com/AI-Zhpp/HAT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 13ab66c0-8b87-46da-89b0-0fb3906806ddCited by top-tier papers17
- Learning Progressive Modality-Shared Transformers for Effective Visible-Infrared Person Re-identificationHu Lu, Xuezhang Zou, Pingping ZhangAAAI 2023 · 183 citations
- Pyramid Spatial-Temporal Aggregation for Video-based Person Re-IdentificationYingquan Wang, Pingping Zhang, Shang Gao, Xia Geng et al.ICCV 2021 · 118 citations
- Cascade Transformers for End-to-End Person SearchRui Yu, Dawei Du, Rodney LaLonde, Daniel Davila et al.CVPR 2022 · 86 citations
- TOP-ReID: Multi-Spectral Object Re-identification with Token PermutationYuhao Wang, Xuehu Liu, Pingping Zhang, Hu Lu et al.AAAI 2024 · 49 citations
- Rotation Invariant Transformer for Recognizing Object in UAVsShuoyi Chen, Mang Ye, Bo DuACM MM 2022 · 43 citations
Builds on16
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang et al.ICCV 2021 · 1,172 citations
- Omni-Scale Feature Learning for Person Re-IdentificationKaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao XiangICCV 2019 · 997 citations
- Rethinking Spatial Dimensions of Vision TransformersByeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun et al.ICCV 2021 · 733 citations
Related papers
- Temporal Correlation Vision Transformer for Video Person Re-IdentificationPengfei Wu, Le Wang, Sanping Zhou, Gang Hua et al.AAAI 2024 · 16 citations
- PSTR: End-to-End One-Step Person Search With TransformersJiale Cao, Yanwei Pang, Rao Muhammad Anwer, Hisham Cholakkal et al.CVPR 2022 · 80 citations
- NFormer: Robust Person Re-identification with Neighbor TransformerHaochen Wang, Jiayi Shen, Yongtuo Liu, Yan Gao et al.CVPR 2022 · 172 citations
- Building Vision Transformers with Hierarchy Aware Feature AggregationYongjie Chen, Hongmin Liu, Haoran Yin, Bin FanICCV 2023 · 6 citations
- Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-IdentificationZhizheng Zhang, Cuiling Lan, Wenjun Zeng, Zhibo ChenCVPR 2020
