Visio-Temporal Attention for Multi-Camera Multi-Target Association
Yu-Jhe Li, Xinshuo Weng, Yan Xu, Kris Kitani
摘要
We address the task of Re-Identification (Re-ID) in multi-target multi-camera (MTMC) tracking where we track multiple pedestrians using multiple overlapping uncalibrated (unknown pose) cameras. Since the videos are temporally synchronized and spatially overlapping, we can see a person from multiple views and associate their trajectory across cameras. In order to find the correct association between pedestrians visible from multiple views during the same time window, we extract a visual feature from a tracklet (sequence of pedestrian images) that encodes its similarity and dissimilarity to all other candidate tracklets. We propose a inter-tracklet (person to person) attention mechanism that learns a representation for a target tracklet while taking into account other tracklets across multiple views. Furthermore, to encode the gait and motion of a person, we introduce second intra-tracklet (person-specific) attention module with position embeddings. This second module employs a transformer encoder to learn a feature from a sequence of features over one tracklet. Experimental results on WILDTRACK and our new dataset ‘ConstructSite’ confirm the superiority of our model over state-of-the-art ReID methods (5% and 10% performance gain respectively) in the context of uncalibrated MTMC tracking. While our model is designed for overlapping cameras, we also obtain state-of-the-art results on two other benchmark datasets (MARS and DukeMTMC) with non-overlapping cameras.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- All-Day Multi-Camera Multi-Target TrackingHuijie Fan, Yu Qiao, Yihao Zhen, Tinghui Zhao 等CVPR 2025
- GMT: Effective Global Framework for Multi-Camera Multi-Target TrackingYihao Zhen, Mingyue Xu, Qiang Wang, Baojie Fan 等CVPR 2026
它引用的顶会 Paper3
- Global-Local Temporal Representations for Video Person Re-IdentificationJianing Li, Shiliang Zhang, Jingdong Wang, Wen Gao 等ICCV 2019 · 被引用 241 次
- Co-Segmentation Inspired Attention Networks for Video-Based Person Re-IdentificationArulkumar Subramaniam, Athira M. Nambiar, Anurag MittalICCV 2019 · 被引用 120 次
- Temporal Knowledge Propagation for Image-to-Video Person Re-IdentificationXinqian Gu, Bingpeng Ma, Hong Chang, Shiguang Shan 等ICCV 2019 · 被引用 65 次
相关 Paper
- Improving Multiple Pedestrian Tracking by Track Management and Occlusion HandlingDaniel Stadler, Jürgen BeyererCVPR 2021
- Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging ScenesKeqi Chen, Vinkle Srivastav, Didier Mutter, Nicolas PadoyCVPR 2025
- Wide-Baseline Multi-Camera Calibration Using Person Re-IdentificationYan Xu, Yu-Jhe Li, Xinshuo Weng, Kris KitaniCVPR 2021
- Cross-Camera Feature Prediction for Intra-Camera Supervised Person Re-identification across Distant ScenesWenhang Ge, Chunyan Pan, Ancong Wu, Hongwei Zheng 等ACM MM 2021 · 被引用 30 次
- Unsupervised Multi-view Pedestrian DetectionMengyin Liu, Chao Zhu, Shiqi Ren, Xu-Cheng YinACM MM 2024 · 被引用 4 次
