Wide-Baseline Multi-Camera Calibration Using Person Re-Identification
Yan Xu, Yu-Jhe Li, Xinshuo Weng, Kris Kitani
摘要
We address the problem of estimating the 3D pose of a network of cameras for large-environment wide-baseline scenarios, e.g., cameras for construction sites, sports stadiums, and public spaces. This task is challenging since detecting and matching the same 3D keypoint observed from two very different camera views is difficult, making standard structure-from-motion (SfM) pipelines inapplicable. In such circumstances, treating people in the scene as "keypoints" and associating them across different camera views can be an alternative method for obtaining correspondences. Based on this intuition, we propose a method that uses ideas from person re-identification (re-ID) for wide-baseline camera calibration. Our method first employs a re-ID method to associate human bounding boxes across cameras, then converts bounding box correspondences to point correspondences, and finally solves for camera pose using multi-view geometry and bundle adjustment. Since our method does not require specialized calibration targets except for visible people, it applies to situations where frequent calibration updates are required. We perform extensive experiments on datasets captured from scenes of different sizes (80m 2 , 350m 2 , 600m 2 ), camera settings (indoor and outdoor), and human activities (walking, playing basketball, construction). Experiment results show that our method achieves similar performance to standard SfM methods relying on manually labeled point correspondences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Virtual Correspondence: Humans as a Cue for Extreme-View GeometryWei-Chiu Ma, Anqi Joyce Yang, Shenlong Wang, Raquel Urtasun 等CVPR 2022 · 被引用 20 次
- Humans as a Calibration Pattern: Dynamic 3D Scene Reconstruction from Unsynchronized and Uncalibrated VideosChangwoon Choi, Jeongjun Kim, Geonho Cha, Minkwan Kim 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper3
- CamNet: Coarse-to-Fine Retrieval for Camera Re-LocalizationMingyu Ding, Zhe Wang, Jiankai Sun, Jianping Shi 等ICCV 2019 · 被引用 163 次
- ELF: Embedded Localisation of Features in Pre-Trained CNNAssia Benbihi, Matthieu Geist, Cédric PradalierICCV 2019 · 被引用 30 次
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
相关 Paper
- Multiview Human Body Reconstruction from Uncalibrated CamerasZhixuan Yu, Linguang Zhang, Yuanlu Xu, Chengcheng Tang 等NeurIPS 2022 · 被引用 24 次
- Reconstructing People, Places, and CamerasLea Müller, Hongsuk Choi, Anthony Zhang, Brent Yi 等CVPR 2025
- Visio-Temporal Attention for Multi-Camera Multi-Target AssociationYu-Jhe Li, Xinshuo Weng, Yan Xu, Kris KitaniICCV 2021 · 被引用 16 次
- MetaPose: Fast 3D Pose from Multiple Views without 3D SupervisionBen Usman, Andrea Tagliasacchi, Kate Saenko, Avneesh SudCVPR 2022 · 被引用 29 次
- RESfM: Robust Deep Equivariant Structure from MotionFadi Khatib, Yoni Kasten, Dror Moran, Meirav Galun 等ICLR 2025
