Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes
Qi Zhang, Jixuan Chen, Zhang Kaiyi, Xinquan Yu, Antoni B. Chan, Hui Huang
Abstract
Multi-view crowd tracking estimates each person's tracking trajectories on the ground of the scene. Recent research works mainly rely on CNNs-based multi-view crowd tracking architectures, and most of them are evaluated and compared on relatively small datasets, such as Wildtrack and MultiviewX. Since these two datasets are collected in small scenes and only contain tens of frames in the evaluation stage, it is difficult for the current methods to be applied to real-world applications where scene size and occlusion are more complicated. In this paper, we propose a Transformer-based multi-view crowd tracking model, MVTrackTrans, which adopts interactions between camera views and the ground plane for enhanced multi-view tracking performance. Besides, for better evaluation, we collect and label two large real-world multi-view tracking datasets, MVCrowdTrack and CityTrack, which contain a much larger scene size over a longer time period. Compared with existing methods on the two large and new datasets, the proposed MVTrackTrans model achieves better performance, demonstrating the advantages of the model design in dealing with large scenes. We believe the proposed datasets and model will push the frontiers of the task to more practical scenarios, and the datasets and code are available at: https://github.com/zqyq/ MVTrackTrans.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ff9e85d-3462-47f0-845d-9bbeed3f37ccBuilds on16
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- MeMOT: Multi-Object Tracking with MemoryJiarui Cai, Mingze Xu, Wei Li, Yuanjun Xiong et al.CVPR 2022 · 216 citations
- Global Tracking TransformersXingyi Zhou, Tianwei Yin, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 180 citations
- Stacked Homography Transformations for Multi-View Pedestrian DetectionLiangchen Song, Jialian Wu, Ming Yang, Qian Zhang et al.ICCV 2021 · 66 citations
Related papers
- MV-TAP: Tracking Any Point in Multi-View VideosJahyeok Koo, Inès Hyeonsu Kim, Mungyeom Kim, Junghyun Park et al.CVPR 2026 · 4 citations
- Cross-View Cross-Scene Multi-View Crowd CountingQi Zhang, Wei Lin, Antoni B. ChanCVPR 2021
- Multi-Person 3D Motion Prediction with Multi-Range TransformersJiashun Wang, Huazhe Xu, Medhini Narasimhan, Xiaolong WangNeurIPS 2021 · 102 citations
- MotionTrack: Learning Robust Short-Term and Long-Term Motions for Multi-Object TrackingZheng Qin, Sanping Zhou, Le Wang, Jinghai Duan et al.CVPR 2023
- DeNoising-MOT: Towards Multiple Object Tracking with Severe OcclusionsTeng Fu, Xiaocong Wang, Haiyang Yu, Ke Niu et al.ACM MM 2023 · 11 citations
