HAT: Hierarchical Aggregation Transformers for Person Re-identification
Guowen Zhang, Pingping Zhang, Jinqing Qi, Huchuan Lu
摘要
Recently, with the advance of deep Convolutional Neural Networks (CNNs), person Re-Identification (Re-ID) has witnessed great success in various applications. However, with limited receptive fields of CNNs, it is still challenging to extract discriminative representations in a global view for persons under non-overlapped cameras. Meanwhile, Transformers demonstrate strong abilities of modeling long-range dependencies for spatial and sequential data. In this work, we take advantages of both CNNs and Transformers, and propose a novel learning framework named Hierarchical Aggregation Transformer (HAT) for image-based person Re-ID with high performance. To achieve this goal, we first propose a Deeply Supervised Aggregation (DSA) to recurrently aggregate hierarchical features from CNN backbones. With multi-granularity supervisions, the DSA can enhance multi-scale features for person retrieval, which is very different from previous methods. Then, we introduce a Transformer-based Feature Calibration (TFC) to integrate low-level detail information as the global prior for high-level semantic information. The proposed TFC is inserted to each level of hierarchical features, resulting in great performance improvements. To our best knowledge, this work is the first to take advantages of both CNNs and Transformers for image-based person Re-ID. Comprehensive experiments on four large-scale Re-ID benchmarks demonstrate that our method shows better results than several state-of-the-art methods. The code is released at https://github.com/AI-Zhpp/HAT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Learning Progressive Modality-Shared Transformers for Effective Visible-Infrared Person Re-identificationHu Lu, Xuezhang Zou, Pingping ZhangAAAI 2023 · 被引用 183 次
- Pyramid Spatial-Temporal Aggregation for Video-based Person Re-IdentificationYingquan Wang, Pingping Zhang, Shang Gao, Xia Geng 等ICCV 2021 · 被引用 118 次
- Cascade Transformers for End-to-End Person SearchRui Yu, Dawei Du, Rodney LaLonde, Daniel Davila 等CVPR 2022 · 被引用 86 次
- TOP-ReID: Multi-Spectral Object Re-identification with Token PermutationYuhao Wang, Xuehu Liu, Pingping Zhang, Hu Lu 等AAAI 2024 · 被引用 49 次
- Rotation Invariant Transformer for Recognizing Object in UAVsShuoyi Chen, Mang Ye, Bo DuACM MM 2022 · 被引用 43 次
它引用的顶会 Paper16
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang 等ICCV 2021 · 被引用 1,172 次
- Omni-Scale Feature Learning for Person Re-IdentificationKaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao XiangICCV 2019 · 被引用 997 次
- Rethinking Spatial Dimensions of Vision TransformersByeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun 等ICCV 2021 · 被引用 733 次
相关 Paper
- Temporal Correlation Vision Transformer for Video Person Re-IdentificationPengfei Wu, Le Wang, Sanping Zhou, Gang Hua 等AAAI 2024 · 被引用 16 次
- PSTR: End-to-End One-Step Person Search With TransformersJiale Cao, Yanwei Pang, Rao Muhammad Anwer, Hisham Cholakkal 等CVPR 2022 · 被引用 80 次
- NFormer: Robust Person Re-identification with Neighbor TransformerHaochen Wang, Jiayi Shen, Yongtuo Liu, Yan Gao 等CVPR 2022 · 被引用 172 次
- Building Vision Transformers with Hierarchy Aware Feature AggregationYongjie Chen, Hongmin Liu, Haoran Yin, Bin FanICCV 2023 · 被引用 6 次
- Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-IdentificationZhizheng Zhang, Cuiling Lan, Wenjun Zeng, Zhibo ChenCVPR 2020
