Pyramid Spatial-Temporal Aggregation for Video-based Person Re-Identification
Yingquan Wang, Pingping Zhang, Shang Gao, Xia Geng, Hu Lu, Dong Wang
Abstract
Video-based person re-identification aims to associate the video clips of the same person across multiple non-overlapping cameras. Spatial-temporal representations can provide richer and complementary information between frames, which are crucial to distinguish the target person when occlusion occurs. This paper proposes a novel Pyramid Spatial-Temporal Aggregation (PSTA) framework to aggregate the frame-level features progressively and fuse the hierarchical temporal features into a final video-level representation. Thus, short-term and long-term temporal information could be well exploited by different hierarchies. Furthermore, a Spatial-Temporal Aggregation Module (STAM) is proposed to enhance the aggregation capability of PSTA. It mainly consists of two novel attention blocks: Spatial Reference Attention (SRA) and Temporal Reference Attention (TRA). SRA explores the spatial correlations within a frame to determine the attention weight of each location. While TRA extends SRA with the correlations between adjacent frames, temporal consistency information can be fully explored to suppress the interference features and strengthen the discriminative ones. Extensive experiments on several challenging benchmarks demonstrate the effectiveness of the proposed PSTA, and our full model reaches 91.5% and 98.3% Rank-1 accuracy on MARS and DukeMTMC-VID benchmarks. The source code is available at https://github.com/WangYQ9/VideoReID-PSTA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30bebeb8-9ed7-413a-99fc-dab156619061Cited by top-tier papers19
- TF-CLIP: Learning Text-Free CLIP for Video-Based Person Re-identificationChenyang Yu, Xuehu Liu, Yingquan Wang, Pingping Zhang et al.AAAI 2024 · 68 citations
- Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-IdentificationPingping Zhang, Yuhao Wang, Yang Liu, Zhengzheng Tu et al.CVPR 2024 · 42 citations
- BigGait: Learning Gait Representation You Want by Large Vision ModelsDingqiang Ye, Chao Fan, Jingzhe Ma, Xiaoming Liu et al.CVPR 2024 · 40 citations
- STPrivacy: Spatio-Temporal Privacy-Preserving Action RecognitionMing Li, Xiangyu Xu, Hehe Fan, Pan Zhou et al.ICCV 2023 · 40 citations
- CLIMB-ReID: A Hybrid CLIP-Mamba Framework for Person Re-IdentificationChenyang Yu, Xuehu Liu, Jiawen Zhu, Yuhao Wang et al.AAAI 2025 · 17 citations
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Mixed High-Order Attention Network for Person Re-IdentificationBinghui Chen, Weihong Deng, Jiani HuICCV 2019 · 392 citations
- Global-Local Temporal Representations for Video Person Re-IdentificationJianing Li, Shiliang Zhang, Jingdong Wang, Wen Gao et al.ICCV 2019 · 241 citations
- HAT: Hierarchical Aggregation Transformers for Person Re-identificationGuowen Zhang, Pingping Zhang, Jinqing Qi, Huchuan LuACM MM 2021 · 159 citations
- Co-Segmentation Inspired Attention Networks for Video-Based Person Re-IdentificationArulkumar Subramaniam, Athira M. Nambiar, Anurag MittalICCV 2019 · 120 citations
Related papers
- ASTA-Net: Adaptive Spatio-Temporal Attention Network for Person Re-Identification in VideosXierong Zhu, Jiawei Liu, Haoze Wu, Meng Wang et al.ACM MM 2020 · 10 citations
- Video-based Person Re-identification with Spatial and Temporal Memory NetworksChanho Eom, Geon Lee, Junghyup Lee, Bumsub HamICCV 2021 · 107 citations
- Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-IdentificationZhizheng Zhang, Cuiling Lan, Wenjun Zeng, Zhibo ChenCVPR 2020
- Temporal Correlation Vision Transformer for Video Person Re-IdentificationPengfei Wu, Le Wang, Sanping Zhou, Gang Hua et al.AAAI 2024 · 16 citations
- Rethinking Temporal Fusion for Video-Based Person Re-Identification on Semantic and Time AspectXinyang Jiang, Yifei Gong, Xiaowei Guo, Qize Yang et al.AAAI 2020 · 21 citations
