Wide-Baseline Multi-Camera Calibration Using Person Re-Identification
Yan Xu, Yu-Jhe Li, Xinshuo Weng, Kris Kitani
Abstract
We address the problem of estimating the 3D pose of a network of cameras for large-environment wide-baseline scenarios, e.g., cameras for construction sites, sports stadiums, and public spaces. This task is challenging since detecting and matching the same 3D keypoint observed from two very different camera views is difficult, making standard structure-from-motion (SfM) pipelines inapplicable. In such circumstances, treating people in the scene as "keypoints" and associating them across different camera views can be an alternative method for obtaining correspondences. Based on this intuition, we propose a method that uses ideas from person re-identification (re-ID) for wide-baseline camera calibration. Our method first employs a re-ID method to associate human bounding boxes across cameras, then converts bounding box correspondences to point correspondences, and finally solves for camera pose using multi-view geometry and bundle adjustment. Since our method does not require specialized calibration targets except for visible people, it applies to situations where frequent calibration updates are required. We perform extensive experiments on datasets captured from scenes of different sizes (80m 2 , 350m 2 , 600m 2 ), camera settings (indoor and outdoor), and human activities (walking, playing basketball, construction). Experiment results show that our method achieves similar performance to standard SfM methods relying on manually labeled point correspondences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d27f6db-e45b-4fe5-9369-20794721a505Cited by top-tier papers2
- Virtual Correspondence: Humans as a Cue for Extreme-View GeometryWei-Chiu Ma, Anqi Joyce Yang, Shenlong Wang, Raquel Urtasun et al.CVPR 2022 · 20 citations
- Humans as a Calibration Pattern: Dynamic 3D Scene Reconstruction from Unsynchronized and Uncalibrated VideosChangwoon Choi, Jeongjun Kim, Geonho Cha, Minkwan Kim et al.ICCV 2025 · 2 citations
Builds on3
- CamNet: Coarse-to-Fine Retrieval for Camera Re-LocalizationMingyu Ding, Zhe Wang, Jiankai Sun, Jianping Shi et al.ICCV 2019 · 163 citations
- ELF: Embedded Localisation of Features in Pre-Trained CNNAssia Benbihi, Matthieu Geist, Cédric PradalierICCV 2019 · 30 citations
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
Related papers
- Multiview Human Body Reconstruction from Uncalibrated CamerasZhixuan Yu, Linguang Zhang, Yuanlu Xu, Chengcheng Tang et al.NeurIPS 2022 · 24 citations
- Reconstructing People, Places, and CamerasLea Müller, Hongsuk Choi, Anthony Zhang, Brent Yi et al.CVPR 2025
- Visio-Temporal Attention for Multi-Camera Multi-Target AssociationYu-Jhe Li, Xinshuo Weng, Yan Xu, Kris KitaniICCV 2021 · 16 citations
- MetaPose: Fast 3D Pose from Multiple Views without 3D SupervisionBen Usman, Andrea Tagliasacchi, Kate Saenko, Avneesh SudCVPR 2022 · 29 citations
- RESfM: Robust Deep Equivariant Structure from MotionFadi Khatib, Yoni Kasten, Dror Moran, Meirav Galun et al.ICLR 2025
