Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging Scenes
Keqi Chen, Vinkle Srivastav, Didier Mutter, Nicolas Padoy
摘要
Multi-view person association is a fundamental step towards multi-view analysis of human activities. Although the person re-identification features have been proven effective, they become unreliable in challenging scenes where persons share similar appearances. Therefore, cross-view geometric constraints are required for a more robust association. However, most existing approaches are either fully-supervised using ground-truth identity labels or require calibrated camera parameters that are hard to obtain. In this work, we investigate the potential of learning from synchronization, and propose a self-supervised uncalibrated multi-view person association approach, Self-MVA, without using any annotations. Specifically, we propose a self-supervised learning framework, consisting of an encoder-decoder model and a self-supervised pretext task, cross-view image synchronization, which aims to distinguish whether two images from different views are captured at the same time. The model encodes each person's unified geometric and appearance features, and we train it by utilizing synchronization labels for supervision after applying Hungarian matching to bridge the gap between instance-wise and image-wise distances. To further reduce the solution space, we propose two types of self-supervised linear constraints: multi-view re-projection and pairwise edge association. Extensive experiments on three challenging public benchmark datasets (WILDTRACK, MVOR, and SOLDIERS) show that our approach achieves state-ofthe-art results, surpassing existing unsupervised and fullysupervised approaches. Code is available at https:// github.com/CAMMA-public/Self-MVA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- Progressive End-to-End Object Detection in Crowded ScenesAnlin Zheng, Yuang Zhang, Xiangyu Zhang, Xiaojuan Qi 等CVPR 2022 · 被引用 82 次
- Stacked Homography Transformations for Multi-View Pedestrian DetectionLiangchen Song, Jialian Wu, Ming Yang, Qian Zhang 等ICCV 2021 · 被引用 66 次
- Shape-aware Multi-Person Pose Estimation from Multi-View ImagesZijian Dong, Jie Song, Xu Chen, Chen Guo 等ICCV 2021 · 被引用 47 次
- Self-supervised Multi-view Multi-Human Association and TrackingYiyang Gan, Ruize Han, Liqiang Yin, Wei Feng 等ACM MM 2021 · 被引用 44 次
相关 Paper
- Self-Supervised Human Pose based Multi-Camera Video SynchronizationLiqiang Yin, Ruize Han, Wei Feng, Song WangACM MM 2022 · 被引用 7 次
- Human Detection and Segmentation via Multi-view ConsensusIsinsu Katircioglu, Helge Rhodin, Jörg Spörri, Mathieu Salzmann 等ICCV 2021 · 被引用 3 次
- CanonPose: Self-Supervised Monocular 3D Human Pose Estimation in the WildBastian Wandt, Marco Rudolph, Petrissa Zell, Helge Rhodin 等CVPR 2021
- Unsupervised Multi-view Pedestrian DetectionMengyin Liu, Chao Zhu, Shiqi Ren, Xu-Cheng YinACM MM 2024 · 被引用 4 次
- Visio-Temporal Attention for Multi-Camera Multi-Target AssociationYu-Jhe Li, Xinshuo Weng, Yan Xu, Kris KitaniICCV 2021 · 被引用 16 次
