CaMuViD: Calibration-Free Multi-View Detection
Amir Etefaghi Daryani, M. Usman Maqbool Bhutta, Byron Hernandez, Henry Medeiros
Abstract
Multi-view object detection in crowded environments presents significant challenges, particularly for occlusion management across multiple camera views. This paper introduces a novel approach that extends conventional multi-view detection to operate directly within each camera's image space. Our method finds objects bounding boxes for images from various perspectives without resorting to a bird's eye view (BEV) representation. Thus, our approach removes the need for camera calibration by leveraging a learnable architecture that facilitates flexible transformations and improves feature fusion across perspectives to increase detection accuracy. Our model achieves Multi-Object Detection Accuracy (MODA) scores of 95.0% and 96.5% on the Wildtrack and MultiviewX datasets, respectively, significantly advancing the state of the art in multi-view detection. Furthermore, it demonstrates robust performance even without ground truth annotations, highlighting its resilience and practicality in real-world applications. These results emphasize the effectiveness of our calibration-free, multi-view object detector.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Unsupervised Multi-View Visual Anomaly Detection via Progressive Homography-Guided AlignmentXintao Chen, Xiaohao Xu, Bozhong Zheng, Yun Liu et al.AAAI 2026 · 1 citation
- Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World ScenesQi Zhang, Jixuan Chen, Zhang Kaiyi, Xinquan Yu et al.CVPR 2026
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- Stacked Homography Transformations for Multi-View Pedestrian DetectionLiangchen Song, Jialian Wu, Ming Yang, Qian Zhang et al.ICCV 2021 · 66 citations
- Multiview Detection with Shadow Transformer (and View-Coherent Data Augmentation)Yunzhong Hou, Liang ZhengACM MM 2021 · 65 citations
- Self-supervised Multi-view Multi-Human Association and TrackingYiyang Gan, Ruize Han, Liqiang Yin, Wei Feng et al.ACM MM 2021 · 44 citations
Related papers
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- From a Bird's Eye View to See: Joint Camera and Subject Registration without the Camera CalibrationZekun Qian, Ruize Han, Wei Feng, Song WangCVPR 2024 · 8 citations
- Towards Generalizable Multi-Camera 3D Object Detection via Perspective RenderingHao Lu, Yunpeng Zhang, Guoqing Wang, Qing Lian et al.AAAI 2025 · 5 citations
- Multiview Human Body Reconstruction from Uncalibrated CamerasZhixuan Yu, Linguang Zhang, Yuanlu Xu, Chengcheng Tang et al.NeurIPS 2022 · 24 citations
- ObjectFusion: Multi-modal 3D Object Detection with Object-Centric FusionQi Cai, Yingwei Pan, Ting Yao, Chong-Wah Ngo et al.ICCV 2023 · 71 citations
