Self-supervised Multi-view Multi-Human Association and Tracking
Yiyang Gan, Ruize Han, Liqiang Yin, Wei Feng, Song Wang
Abstract
Multi-view Multi-human association and tracking (MvMHAT) aims to track a group of people over time in each view, as well as to identify the same person across different views at the same time. This is a relatively new problem but is very important for multi-person scene video surveillance. Different from previous multiple object tracking (MOT) and multi-target multi-camera tracking (MTMCT) tasks, which only consider the over-time human association, MvMHAT requires to jointly achieve both cross-view and over-time data association. In this paper, we model this problem with a self-supervised learning framework and leverage an end-to-end network to tackle it. Specifically, we propose a spatial-temporal association network with two designed self-supervised learning losses, including a symmetric-similarity loss and a transitive-similarity loss, at each time to associate the multiple humans over time and across views. Besides, to promote the research on MvMHAT, we build a new large-scale benchmark for the training and testing of different algorithms. Extensive experiments on the proposed benchmark verify the effectiveness of our method. We have released the benchmark and code to the public.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43971e31-fb92-4dc7-80f0-5110d954d82eCited by top-tier papers12
- Safe Multi-View Deep ClassificationWei Liu, Yufei Chen, Xiaodong Yue, Changqing Zhang et al.AAAI 2023 · 27 citations
- Mimicking the Annotation Process for Recognizing the Micro ExpressionsBo-Kai Ruan, Ling Lo, Hong-Han Shuai, Wen-Huang ChengACM MM 2022 · 16 citations
- Cross-View Referring Multi-Object TrackingSijia Chen, En Yu, Wenbing TaoAAAI 2025 · 15 citations
- Connecting the Complementary-view Videos: Joint Camera Identification and Subject AssociationRuize Han, Yiyang Gan, Jiacheng Li, Feifan Wang et al.CVPR 2022 · 12 citations
- Multi-View Pedestrian Occupancy Prediction with a Novel Synthetic DatasetSithu Aung, Min-Cheol Sagong, Junghyun ChoAAAI 2025 · 5 citations
Builds on9
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- Spatial-Temporal Relation Networks for Multi-Object TrackingJiarui Xu, Yue Cao, Zheng Zhang, Han HuICCV 2019 · 260 citations
- FAMNet: Joint Learning of Feature, Affinity and Multi-Dimensional Assignment for Online Multiple Object TrackingPeng Chu, Haibin LingICCV 2019 · 229 citations
- Unsupervised Graph Association for Person Re-IdentificationJinlin Wu, Hao Liu, Yang Yang, Zhen Lei et al.ICCV 2019 · 116 citations
- Complementary-View Multiple Human TrackingRuize Han, Wei Feng, Jiewen Zhao, Zicheng Niu et al.AAAI 2020 · 36 citations
Related papers
- Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging ScenesKeqi Chen, Vinkle Srivastav, Didier Mutter, Nicolas PadoyCVPR 2025
- Self-Supervised Human Pose based Multi-Camera Video SynchronizationLiqiang Yin, Ruize Han, Wei Feng, Song WangACM MM 2022 · 7 citations
- MVTrajecter: Multi-View Pedestrian Tracking With Trajectory Motion Cost and Trajectory Appearance CostTaiga Yamane, Ryo Masumura, Satoshi Suzuki, Shota OrihashiICCV 2025 · 2 citations
- Cross-View Tracking for Multi-Human 3D Pose Estimation at Over 100 FPSLong Chen, Haizhou Ai, Rui Chen, Zijie Zhuang et al.CVPR 2020
- Multi-View Domain Adaptive Object Detection on Camera NetworksYan Lu, Zhun Zhong, Yuanchao ShuAAAI 2023 · 4 citations
