From a Bird's Eye View to See: Joint Camera and Subject Registration without the Camera Calibration
Zekun Qian, Ruize Han, Wei Feng, Song Wang
摘要
We tackle a new problem of multi-view camera and sub-ject registration in the bird’ s eye view (BEV) without pregiven camera calibration, which promotes the multi-view subject registration problem to a new calibration-free stage. This greatly alleviates the limitation in many practical applications. However, this is a very challenging problem since its only input is several RGB images from different first-person views (FPVs), without the BEV image and the calibration of the FPVs, while the output is a unified plane aggregated from all views with the positions and orientations of both the subjects and cameras in a BEV. For this purpose, we propose an end-to-end framework solving cam-era and subject registration together by taking advantage of their mutual dependence, whose main idea is as below: i) creating a subject view-transform module (VTM) to project each pedestrian from FPV to a virtual BEV, ii) deriving a multi-view geometry-based spatial alignment module (SAM) to estimate the relative camera pose in a unified BEV, iii) selecting and refining the subject and camera registration results within the unified BEV. We collect a new large-scale synthetic dataset with rich annotations for training and evaluation. Additionally, we also collect a real dataset for cross-domain evaluation. The experimental results show the remarkable effectiveness of our method. The code and proposed datasets are available at BEVSee.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large ModelsJunkang Liu, Fanhua Shang, Hongying Liu, Yuxuan Tian 等AAAI 2026 · 被引用 12 次
- MVTrajecter: Multi-View Pedestrian Tracking With Trajectory Motion Cost and Trajectory Appearance CostTaiga Yamane, Ryo Masumura, Satoshi Suzuki, Shota OrihashiICCV 2025 · 被引用 2 次
它引用的顶会 Paper20
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- PolarFormer: Multi-Camera 3D Object Detection with Polar TransformerYanqin Jiang, Li Zhang, Zhenwei Miao, Xiatian Zhu 等AAAI 2023 · 被引用 240 次
- MonoLoco: Monocular 3D Pedestrian Localization and Uncertainty EstimationLorenzo Bertoni, Sven Kreiss, Alexandre AlahiICCV 2019 · 被引用 125 次
- Correlation Verification for Image RetrievalSeongwon Lee, Hongje Seong, Suhyeon Lee, Euntai KimCVPR 2022 · 被引用 79 次
- Stacked Homography Transformations for Multi-View Pedestrian DetectionLiangchen Song, Jialian Wu, Ming Yang, Qian Zhang 等ICCV 2021 · 被引用 66 次
相关 Paper
- CaMuViD: Calibration-Free Multi-View DetectionAmir Etefaghi Daryani, M. Usman Maqbool Bhutta, Byron Hernandez, Henry MedeirosCVPR 2025
- VGA: Empowering Aerial-Ground Localization by Visual Geometry AlignmentTao Jun Lin, Yujiao Shi, Hongdong LiCVPR 2026
- Putting People in their Place: Monocular Regression of 3D People in DepthYu Sun, Wu Liu, Qian Bao, Yili Fu 等CVPR 2022 · 被引用 152 次
- Multiview Human Body Reconstruction from Uncalibrated CamerasZhixuan Yu, Linguang Zhang, Yuanlu Xu, Chengcheng Tang 等NeurIPS 2022 · 被引用 24 次
- Unsupervised Multi-view Pedestrian DetectionMengyin Liu, Chao Zhu, Shiqi Ren, Xu-Cheng YinACM MM 2024 · 被引用 4 次
