Pyramid Spatial-Temporal Aggregation for Video-based Person Re-Identification
Yingquan Wang, Pingping Zhang, Shang Gao, Xia Geng, Hu Lu, Dong Wang
摘要
Video-based person re-identification aims to associate the video clips of the same person across multiple non-overlapping cameras. Spatial-temporal representations can provide richer and complementary information between frames, which are crucial to distinguish the target person when occlusion occurs. This paper proposes a novel Pyramid Spatial-Temporal Aggregation (PSTA) framework to aggregate the frame-level features progressively and fuse the hierarchical temporal features into a final video-level representation. Thus, short-term and long-term temporal information could be well exploited by different hierarchies. Furthermore, a Spatial-Temporal Aggregation Module (STAM) is proposed to enhance the aggregation capability of PSTA. It mainly consists of two novel attention blocks: Spatial Reference Attention (SRA) and Temporal Reference Attention (TRA). SRA explores the spatial correlations within a frame to determine the attention weight of each location. While TRA extends SRA with the correlations between adjacent frames, temporal consistency information can be fully explored to suppress the interference features and strengthen the discriminative ones. Extensive experiments on several challenging benchmarks demonstrate the effectiveness of the proposed PSTA, and our full model reaches 91.5% and 98.3% Rank-1 accuracy on MARS and DukeMTMC-VID benchmarks. The source code is available at https://github.com/WangYQ9/VideoReID-PSTA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- TF-CLIP: Learning Text-Free CLIP for Video-Based Person Re-identificationChenyang Yu, Xuehu Liu, Yingquan Wang, Pingping Zhang 等AAAI 2024 · 被引用 68 次
- Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-IdentificationPingping Zhang, Yuhao Wang, Yang Liu, Zhengzheng Tu 等CVPR 2024 · 被引用 42 次
- BigGait: Learning Gait Representation You Want by Large Vision ModelsDingqiang Ye, Chao Fan, Jingzhe Ma, Xiaoming Liu 等CVPR 2024 · 被引用 40 次
- STPrivacy: Spatio-Temporal Privacy-Preserving Action RecognitionMing Li, Xiangyu Xu, Hehe Fan, Pan Zhou 等ICCV 2023 · 被引用 40 次
- CLIMB-ReID: A Hybrid CLIP-Mamba Framework for Person Re-IdentificationChenyang Yu, Xuehu Liu, Jiawen Zhu, Yuhao Wang 等AAAI 2025 · 被引用 17 次
它引用的顶会 Paper12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Mixed High-Order Attention Network for Person Re-IdentificationBinghui Chen, Weihong Deng, Jiani HuICCV 2019 · 被引用 392 次
- Global-Local Temporal Representations for Video Person Re-IdentificationJianing Li, Shiliang Zhang, Jingdong Wang, Wen Gao 等ICCV 2019 · 被引用 241 次
- HAT: Hierarchical Aggregation Transformers for Person Re-identificationGuowen Zhang, Pingping Zhang, Jinqing Qi, Huchuan LuACM MM 2021 · 被引用 159 次
- Co-Segmentation Inspired Attention Networks for Video-Based Person Re-IdentificationArulkumar Subramaniam, Athira M. Nambiar, Anurag MittalICCV 2019 · 被引用 120 次
相关 Paper
- ASTA-Net: Adaptive Spatio-Temporal Attention Network for Person Re-Identification in VideosXierong Zhu, Jiawei Liu, Haoze Wu, Meng Wang 等ACM MM 2020 · 被引用 10 次
- Video-based Person Re-identification with Spatial and Temporal Memory NetworksChanho Eom, Geon Lee, Junghyup Lee, Bumsub HamICCV 2021 · 被引用 107 次
- Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-IdentificationZhizheng Zhang, Cuiling Lan, Wenjun Zeng, Zhibo ChenCVPR 2020
- Temporal Correlation Vision Transformer for Video Person Re-IdentificationPengfei Wu, Le Wang, Sanping Zhou, Gang Hua 等AAAI 2024 · 被引用 16 次
- Rethinking Temporal Fusion for Video-Based Person Re-Identification on Semantic and Time AspectXinyang Jiang, Yifei Gong, Xiaowei Guo, Qize Yang 等AAAI 2020 · 被引用 21 次
