Visio-Temporal Attention for Multi-Camera Multi-Target Association
Yu-Jhe Li, Xinshuo Weng, Yan Xu, Kris Kitani
Abstract
We address the task of Re-Identification (Re-ID) in multi-target multi-camera (MTMC) tracking where we track multiple pedestrians using multiple overlapping uncalibrated (unknown pose) cameras. Since the videos are temporally synchronized and spatially overlapping, we can see a person from multiple views and associate their trajectory across cameras. In order to find the correct association between pedestrians visible from multiple views during the same time window, we extract a visual feature from a tracklet (sequence of pedestrian images) that encodes its similarity and dissimilarity to all other candidate tracklets. We propose a inter-tracklet (person to person) attention mechanism that learns a representation for a target tracklet while taking into account other tracklets across multiple views. Furthermore, to encode the gait and motion of a person, we introduce second intra-tracklet (person-specific) attention module with position embeddings. This second module employs a transformer encoder to learn a feature from a sequence of features over one tracklet. Experimental results on WILDTRACK and our new dataset ‘ConstructSite’ confirm the superiority of our model over state-of-the-art ReID methods (5% and 10% performance gain respectively) in the context of uncalibrated MTMC tracking. While our model is designed for overlapping cameras, we also obtain state-of-the-art results on two other benchmark datasets (MARS and DukeMTMC) with non-overlapping cameras.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6bd2a88f-22fc-46b0-8a98-9f70765754f8Cited by top-tier papers2
- All-Day Multi-Camera Multi-Target TrackingHuijie Fan, Yu Qiao, Yihao Zhen, Tinghui Zhao et al.CVPR 2025
- GMT: Effective Global Framework for Multi-Camera Multi-Target TrackingYihao Zhen, Mingyue Xu, Qiang Wang, Baojie Fan et al.CVPR 2026
Builds on3
- Global-Local Temporal Representations for Video Person Re-IdentificationJianing Li, Shiliang Zhang, Jingdong Wang, Wen Gao et al.ICCV 2019 · 241 citations
- Co-Segmentation Inspired Attention Networks for Video-Based Person Re-IdentificationArulkumar Subramaniam, Athira M. Nambiar, Anurag MittalICCV 2019 · 120 citations
- Temporal Knowledge Propagation for Image-to-Video Person Re-IdentificationXinqian Gu, Bingpeng Ma, Hong Chang, Shiguang Shan et al.ICCV 2019 · 65 citations
Related papers
- Improving Multiple Pedestrian Tracking by Track Management and Occlusion HandlingDaniel Stadler, Jürgen BeyererCVPR 2021
- Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging ScenesKeqi Chen, Vinkle Srivastav, Didier Mutter, Nicolas PadoyCVPR 2025
- Wide-Baseline Multi-Camera Calibration Using Person Re-IdentificationYan Xu, Yu-Jhe Li, Xinshuo Weng, Kris KitaniCVPR 2021
- Cross-Camera Feature Prediction for Intra-Camera Supervised Person Re-identification across Distant ScenesWenhang Ge, Chunyan Pan, Ancong Wu, Hongwei Zheng et al.ACM MM 2021 · 30 citations
- Unsupervised Multi-view Pedestrian DetectionMengyin Liu, Chao Zhu, Shiqi Ren, Xu-Cheng YinACM MM 2024 · 4 citations
