GMT: Effective Global Framework for Multi-Camera Multi-Target Tracking
Yihao Zhen, Mingyue Xu, Qiang Wang, Baojie Fan, Jiahua Dong, Tinghui Zhao, Huijie Fan
Abstract
Multi-Camera Multi-Target (MCMT) tracking aims to locate and associate the same targets across multiple camera views. Existing methods typically adopt a two-stage framework, involving single-camera tracking followed by inter-camera tracking. However, in this paradigm, multiview information is used only to recover missed matches in the first stage, providing a limited contribution to overall tracking. To address this issue, we propose GMT, a global MCMT tracking framework that jointly exploits intra-view and inter-view cues for tracking. Specifically, instead of assigning trajectories independently for each view, we integrate the same historical targets across different views as global trajectories, thereby reformulating the two-stage tracking as a unified global-level trajectory-target association process. We introduce a Cross-View Feature Consistency Enhancement (CFCE) module to align visual and spatial features across views, providing a consistent feature space for global trajectory modeling. With these aligned features, the Global Trajectory Association (GTA) module associates new detections with existing global trajectories, enabling direct use of multi-view information. Compared to the two-stage framework, GMT achieves significant improvements on existing datasets, with gains of up to 21.3 percent in CVMA and 17.2 percent in CVIDF1. Furthermore, we introduce VisionTrack, a high-quality, largescale MCMT dataset providing significantly greater diversity than existing datasets. Our code and dataset will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0e067a2-ef80-4be8-9421-d7d17db6ea8eBuilds on13
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- Bayesian Loss for Crowd Count Estimation With Point SupervisionZhiheng Ma, Xing Wei, Xiaopeng Hong, Yihong GongICCV 2019 · 612 citations
- Infrared-Visible Cross-Modal Person Re-Identification with an X ModalityDiangang Li, Xing Wei, Xiaopeng Hong, Yihong GongAAAI 2020 · 419 citations
- Global Tracking TransformersXingyi Zhou, Tianwei Yin, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 180 citations
- LMGP: Lifted Multicut Meets Geometry Projections for Multi-Camera Multi-Object TrackingDuy M. H. Nguyen, Roberto Henschel, Bodo Rosenhahn, Daniel Sonntag et al.CVPR 2022 · 49 citations
Related papers
- Traffic-Aware Multi-Camera Tracking of Vehicles Based on ReID and Camera Link ModelHung-Min Hsu, Yizhou Wang, Jenq-Neng HwangACM MM 2020 · 46 citations
- ReST: A Reconfigurable Spatial-Temporal Graph Model for Multi-Camera Multi-Object TrackingCheng-Che Cheng, Min-Xuan Qiu, Chen-Kuo Chiang, Shang-Hong LaiICCV 2023 · 39 citations
- Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured VideosJianbo Ma, Hui Luo, Qi Chen, Yuankai Qi et al.AAAI 2026 · 2 citations
- Visio-Temporal Attention for Multi-Camera Multi-Target AssociationYu-Jhe Li, Xinshuo Weng, Yan Xu, Kris KitaniICCV 2021 · 16 citations
- MITracker: Multi-View Integration for Visual Object TrackingMengjie Xu, Yitao Zhu, Haotian Jiang, Jiaming Li et al.CVPR 2025
