Simultaneous Multi-View Instance Detection With Learned Geometric Soft-Constraints
Ahmed Samy Nassar, Sébastien Lefèvre, Jan Dirk Wegner
Abstract
We propose to jointly learn multi-view geometry and warping between views of the same object instances for robust cross-view object detection. What makes multi-view object instance detection difficult are strong changes in viewpoint, lighting conditions, high similarity of neighbouring objects, and strong variability in scale. By turning object detection and instance re-identification in different views into a joint learning task, we are able to incorporate both image appearance and geometric soft constraints into a single, multi-view detection process that is learnable end-to-end. We validate our method on a new, large data set of street-level panoramas of urban objects and show superior performance compared to various baselines. Our contribution is threefold: a large-scale, publicly available data set for multi-view instance detection and re-identification; an annotation tool custom-tailored for multi-view instance detection; and a novel, holistic multi-view instance detection and re-identification method that jointly models geometry and appearance across views.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Stacked Homography Transformations for Multi-View Pedestrian DetectionLiangchen Song, Jialian Wu, Ming Yang, Qian Zhang et al.ICCV 2021 · 66 citations
- The Auto Arborist Dataset: A Large-Scale Benchmark for Multiview Urban Forest Monitoring Under Domain ShiftSara Beery, Guanhang Wu, Trevor Edwards, Filip Pavetic et al.CVPR 2022 · 53 citations
- SAIL-VOS 3D: A Synthetic Dataset and Baselines for Object Detection and 3D Mesh Reconstruction From Video DataYuan-Ting Hu, Jiahong Wang, Raymond A. Yeh, Alexander G. SchwingCVPR 2021
Related papers
- Geometry-Aware Satellite-to-Ground Image Synthesis for Urban AreasXiaohu Lu, Zuoyue Li, Zhaopeng Cui, Martin R. Oswald et al.CVPR 2020
- Viewpoint Equivariance for Multi-View 3D Object DetectionDian Chen, Jie Li, Vitor Guizilini, Rares Ambrus et al.CVPR 2023
- Detection Based Part-level Articulated Object Reconstruction from Single RGBD ImageYuki Kawana, Tatsuya HaradaNeurIPS 2023 · 20 citations
- Norm-Aware Embedding for Efficient Person SearchDi Chen, Shanshan Zhang, Jian Yang, Bernt SchieleCVPR 2020
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object DetectionYung-Hsu Yang, Luigi Piccinelli, Mattia Segù, Siyuan Li et al.ICCV 2025 · 2 citations
