Dense Interaction Learning for Video-based Person Re-identification
Tianyu He, Xin Jin, Xu Shen, Jianqiang Huang, Zhibo Chen, Xian-Sheng Hua
摘要
Video-based person re-identification (re-ID) aims at matching the same person across video clips. Efficiently exploiting multi-scale fine-grained features while building the structural interaction among them is pivotal for its success. In this paper, we propose a hybrid framework, Dense Interaction Learning (DenseIL), that takes the principal advantages of both CNN-based and Attention-based architectures to tackle video-based person re-ID difficulties. DenseIL contains a CNN encoder and a Dense Interaction (DI) decoder. The CNN encoder is responsible for efficiently extracting discriminative spatial features while the DI decoder is designed to densely model spatial-temporal inherent interaction across frames. Different from previous works, we additionally let the DI decoder densely attends to intermediate fine-grained CNN features and that naturally yields multi-grained spatial-temporal representation for each video clip. Moreover, we introduce Spatio-TEmporal Positional Embedding (STEP-Emb) into the DI decoder to investigate the positional relation among the spatial-temporal inputs. Our experiments consistently and significantly outperform all the state-of-the-art methods on multiple standard video-based person re-ID datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Salient-to-Broad Transition for Video Person Re-identificationShutao Bai, Bingpeng Ma, Hong Chang, Rui Huang 等CVPR 2022 · 被引用 71 次
- TF-CLIP: Learning Text-Free CLIP for Video-Based Person Re-identificationChenyang Yu, Xuehu Liu, Yingquan Wang, Pingping Zhang 等AAAI 2024 · 被引用 68 次
- Multi-Prompts Learning with Cross-Modal Alignment for Attribute-Based Person Re-identificationYajing Zhai, Yawen Zeng, Zhiyong Huang, Zheng Qin 等AAAI 2024 · 被引用 40 次
- CLIMB-ReID: A Hybrid CLIP-Mamba Framework for Person Re-IdentificationChenyang Yu, Xuehu Liu, Jiawen Zhu, Yuhao Wang 等AAAI 2025 · 被引用 17 次
- Temporal Correlation Vision Transformer for Video Person Re-IdentificationPengfei Wu, Le Wang, Sanping Zhou, Gang Hua 等AAAI 2024 · 被引用 16 次
它引用的顶会 Paper19
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li 等AAAI 2020 · 被引用 4,134 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
相关 Paper
- Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-IdentificationZhizheng Zhang, Cuiling Lan, Wenjun Zeng, Zhibo ChenCVPR 2020
- Rethinking Temporal Fusion for Video-Based Person Re-Identification on Semantic and Time AspectXinyang Jiang, Yifei Gong, Xiaowei Guo, Qize Yang 等AAAI 2020 · 被引用 21 次
- Watching You: Global-Guided Reciprocal Learning for Video-Based Person Re-IdentificationXuehu Liu, Pingping Zhang, Chenyang Yu, Huchuan Lu 等CVPR 2021
- Spatial-Temporal Correlation and Topology Learning for Person Re-Identification in VideosJiawei Liu, Zheng-Jun Zha, Wei Wu, Kecheng Zheng 等CVPR 2021
- Pyramid Spatial-Temporal Aggregation for Video-based Person Re-IdentificationYingquan Wang, Pingping Zhang, Shang Gao, Xia Geng 等ICCV 2021 · 被引用 118 次
