Global-Local Temporal Representations for Video Person Re-Identification
Jianing Li, Shiliang Zhang, Jingdong Wang, Wen Gao, Qi Tian
Abstract
This paper proposes the Global-Local Temporal Representation (GLTR) to exploit the multi-scale temporal cues in video sequences for video person Re-Identification (ReID). GLTR is constructed by first modeling the short-term temporal cues among adjacent frames, then capturing the long-term relations among inconsecutive frames. Specifically, the short-term temporal cues are modeled by parallel dilated convolutions with different temporal dilation rates to represent the motion and appearance of pedestrian. The long-term relations are captured by a temporal self-attention model to alleviate the occlusions and noises in video sequences. The short and long-term temporal cues are aggregated as the final GLTR by a simple single-stream CNN. GLTR shows substantial superiority to existing features learned with body part cues or metric learning on four widely-used video ReID datasets. For instance, it achieves Rank-1 Accuracy of 87.02% on MARS dataset without re-ranking, better than current state-of-the art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fed3fc2e-5349-450c-87c6-69ea9f75e4efCited by top-tier papers36
- Clothes-Changing Person Re-identification with RGB Modality OnlyXinqian Gu, Hong Chang, Bingpeng Ma, Shutao Bai et al.CVPR 2022 · 226 citations
- Pyramid Spatial-Temporal Aggregation for Video-based Person Re-IdentificationYingquan Wang, Pingping Zhang, Shang Gao, Xia Geng et al.ICCV 2021 · 118 citations
- Video-based Person Re-identification with Spatial and Temporal Memory NetworksChanho Eom, Geon Lee, Junghyup Lee, Bumsub HamICCV 2021 · 107 citations
- Gait Recognition in the Wild: A BenchmarkICCV 2021 · 102 citations
- Spatio-Temporal Representation Factorization for Video-based Person Re-IdentificationAbhishek Aich, Meng Zheng, Srikrishna Karanam, Terrence Chen et al.ICCV 2021 · 86 citations
Builds on1
Related papers
- ASTA-Net: Adaptive Spatio-Temporal Attention Network for Person Re-Identification in VideosXierong Zhu, Jiawei Liu, Haoze Wu, Meng Wang et al.ACM MM 2020 · 10 citations
- Spatial-Temporal Correlation and Topology Learning for Person Re-Identification in VideosJiawei Liu, Zheng-Jun Zha, Wei Wu, Kecheng Zheng et al.CVPR 2021
- Learning Multi-Granular Hypergraphs for Video-Based Person Re-IdentificationYichao Yan, Jie Qin, Jiaxin Chen, Li Liu et al.CVPR 2020
- Spatial-Temporal Graph Convolutional Network for Video-Based Person Re-IdentificationJinrui Yang, Wei-Shi Zheng, Qize Yang, Ying-Cong Chen et al.CVPR 2020
- Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-IdentificationZhizheng Zhang, Cuiling Lan, Wenjun Zeng, Zhibo ChenCVPR 2020
