Rethinking Temporal Fusion for Video-Based Person Re-Identification on Semantic and Time Aspect
Xinyang Jiang, Yifei Gong, Xiaowei Guo, Qize Yang, Feiyue Huang, Wei-Shi Zheng, Feng Zheng, Xing Sun
Abstract
Recently, the research interest of person re-identification (ReID) has gradually turned to video-based methods, which acquire a person representation by aggregating frame features of an entire video. However, existing video-based ReID methods do not consider the semantic difference brought by the outputs of different network stages, which potentially compromises the information richness of the person features. Furthermore, traditional methods ignore important relationship among frames, which causes information redundancy in fusion along the time axis. To address these issues, we propose a novel general temporal fusion framework to aggregate frame features on both semantic aspect and time aspect. As for the semantic aspect, a multi-stage fusion network is explored to fuse richer frame features at multiple semantic levels, which can effectively reduce the information loss caused by the traditional single-stage fusion. While, for the time axis, the existing intra-frame attention method is improved by adding a novel inter-frame attention module, which effectively reduces the information redundancy in temporal fusion by taking the relationship among frames into consideration. The experimental results show that our approach can effectively improve the video-based re-identification accuracy, achieving the state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6429538d-6190-4e80-8cb2-c1c20d7b021dCited by top-tier papers4
- Pyramid Spatial-Temporal Aggregation for Video-based Person Re-IdentificationYingquan Wang, Pingping Zhang, Shang Gao, Xia Geng et al.ICCV 2021 · 118 citations
- DisenQ: Disentangling Q-Former for Activity-BiometricsShehreen Azad, Yogesh Singh RawatICCV 2025 · 4 citations
- Spatial-Temporal Correlation and Topology Learning for Person Re-Identification in VideosJiawei Liu, Zheng-Jun Zha, Wei Wu, Kecheng Zheng et al.CVPR 2021
- Activity-Biometrics: Person Identification from Daily ActivitiesShehreen Azad, Yogesh Singh RawatCVPR 2024
Related papers
- Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-IdentificationZhizheng Zhang, Cuiling Lan, Wenjun Zeng, Zhibo ChenCVPR 2020
- Relation-Guided Spatial Attention and Temporal Refinement for Video-Based Person Re-IdentificationXingze Li, Wengang Zhou, Yun Zhou, Houqiang LiAAAI 2020 · 33 citations
- Viewing from Frequency Domain: A DCT-based Information Enhancement Network for Video Person Re-IdentificationLiangchen Liu, Xi Yang, Nannan Wang, Xinbo GaoACM MM 2021 · 11 citations
- ASTA-Net: Adaptive Spatio-Temporal Attention Network for Person Re-Identification in VideosXierong Zhu, Jiawei Liu, Haoze Wu, Meng Wang et al.ACM MM 2020 · 10 citations
- Dense Interaction Learning for Video-based Person Re-identificationTianyu He, Xin Jin, Xu Shen, Jianqiang Huang et al.ICCV 2021 · 73 citations
