Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging Scenes
Keqi Chen, Vinkle Srivastav, Didier Mutter, Nicolas Padoy
Abstract
Multi-view person association is a fundamental step towards multi-view analysis of human activities. Although the person re-identification features have been proven effective, they become unreliable in challenging scenes where persons share similar appearances. Therefore, cross-view geometric constraints are required for a more robust association. However, most existing approaches are either fully-supervised using ground-truth identity labels or require calibrated camera parameters that are hard to obtain. In this work, we investigate the potential of learning from synchronization, and propose a self-supervised uncalibrated multi-view person association approach, Self-MVA, without using any annotations. Specifically, we propose a self-supervised learning framework, consisting of an encoder-decoder model and a self-supervised pretext task, cross-view image synchronization, which aims to distinguish whether two images from different views are captured at the same time. The model encodes each person's unified geometric and appearance features, and we train it by utilizing synchronization labels for supervision after applying Hungarian matching to bridge the gap between instance-wise and image-wise distances. To further reduce the solution space, we propose two types of self-supervised linear constraints: multi-view re-projection and pairwise edge association. Extensive experiments on three challenging public benchmark datasets (WILDTRACK, MVOR, and SOLDIERS) show that our approach achieves state-ofthe-art results, surpassing existing unsupervised and fullysupervised approaches. Code is available at https:// github.com/CAMMA-public/Self-MVA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0850c248-a32e-4028-b74c-66524b8b89cdCited by top-tier papers1
Ask how each one uses itBuilds on11
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- Progressive End-to-End Object Detection in Crowded ScenesAnlin Zheng, Yuang Zhang, Xiangyu Zhang, Xiaojuan Qi et al.CVPR 2022 · 82 citations
- Stacked Homography Transformations for Multi-View Pedestrian DetectionLiangchen Song, Jialian Wu, Ming Yang, Qian Zhang et al.ICCV 2021 · 66 citations
- Shape-aware Multi-Person Pose Estimation from Multi-View ImagesZijian Dong, Jie Song, Xu Chen, Chen Guo et al.ICCV 2021 · 47 citations
- Self-supervised Multi-view Multi-Human Association and TrackingYiyang Gan, Ruize Han, Liqiang Yin, Wei Feng et al.ACM MM 2021 · 44 citations
Related papers
- Self-Supervised Human Pose based Multi-Camera Video SynchronizationLiqiang Yin, Ruize Han, Wei Feng, Song WangACM MM 2022 · 7 citations
- Human Detection and Segmentation via Multi-view ConsensusIsinsu Katircioglu, Helge Rhodin, Jörg Spörri, Mathieu Salzmann et al.ICCV 2021 · 3 citations
- CanonPose: Self-Supervised Monocular 3D Human Pose Estimation in the WildBastian Wandt, Marco Rudolph, Petrissa Zell, Helge Rhodin et al.CVPR 2021
- Unsupervised Multi-view Pedestrian DetectionMengyin Liu, Chao Zhu, Shiqi Ren, Xu-Cheng YinACM MM 2024 · 4 citations
- Visio-Temporal Attention for Multi-Camera Multi-Target AssociationYu-Jhe Li, Xinshuo Weng, Yan Xu, Kris KitaniICCV 2021 · 16 citations
