'The Pedestrian next to the Lamppost" Adaptive Object Graphs for Better Instantaneous Mapping
Avishkar Saha, Oscar Mendez, Chris Russell, Richard Bowden
Abstract
Estimating a semantically segmented bird's-eye-view (BEV) map from a single image has become a popular technique for autonomous control and navigation. However, they show an increase in localization error with distance from the camera. While such an increase in error is entirely expected -localization is harder at distance - much of the drop in performance can be attributed to the cues used by current texture-based models, in particular, they make heavy use of object-ground intersections (such as shadows) [10], which become increasingly sparse and uncertain for distant objects. In this work, we address these shortcomings in BEV-mapping by learning the spatial relationship between objects in a scene. We propose a graph neural network which predicts BEV objects from a monocular image by spatially reasoning about an object within the context of other objects. Our approach sets a new state-of-the-art in BEV estimation from monocular images across three large-scale datasets, including a 50% relative improvement for objects on nuScenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a03078ea-5957-4bea-b05a-2352bcf7c326Builds on9
- Rethinking Graph Transformers with Spectral AttentionDevin Kreuzer, Dominique Beaini, William L. Hamilton, Vincent Létourneau et al.NeurIPS 2021 · 854 citations
- Disentangling Monocular 3D Object DetectionAndrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Manuel Lopez-Antequera et al.ICCV 2019 · 504 citations
- Directional Graph NetworksDominique Beaini, Saro Passaro, Vincent Létourneau, William L. Hamilton et al.ICML 2021 · 216 citations
- How Do Neural Networks See Depth in Single Images?Tom van Dijk, Guido de CroonICCV 2019 · 210 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
Related papers
- Predicting Semantic Map Representations From Images Using Pyramid Occupancy NetworksThomas Roddick, Roberto CipollaCVPR 2020
- OccluBEV: Occlusion Aware Spatiotemporal Modeling for Multi-view 3D Object DetectionZiteng Wen, Hai Xu, Chenyu Liu, Tao Guo et al.ACM MM 2023 · 5 citations
- BEV-CAR: Enhancing Monocular Bird's Eye View Segmentation with Context-Aware RasterizationYixin Xiong, Ke Wang, Tongtong Cheng, Chunhui Liu et al.CVPR 2026
- VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector QuantizationYiwei Zhang, Jin Gao, Fudong Ge, Guan Luo et al.NeurIPS 2024 · 3 citations
- Instance-Aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting LearningYang Jiao, Zequn Jie, Shaoxiang Chen, Lechao Cheng et al.AAAI 2024 · 13 citations
