Weakly But Deeply Supervised Occlusion-Reasoned Parametric Road Layouts
Buyu Liu, Bingbing Zhuang, Manmohan Chandraker
Abstract
We propose an end-to-end network that takes a single perspective RGB image of a complex road scene as input, to produce occlusion-reasoned layouts in perspective space as well as a parametric bird's-eye-view (BEV) space. In contrast to prior works that require dense supervision such as semantic labels in perspective view, our method only requires human annotations for parametric attributes that are cheaper and less ambiguous to obtain. To solve this challenging task, our design is comprised of modules that incorporate inductive biases to learn occlusion-reasoning, geometric transformation and semantic abstraction, where each module may be supervised by appropriately transforming the parametric annotations. We demonstrate how our design choices and proposed deep supervision help achieve meaningful representations and accurate predictions. We validate our approach on two public datasets, KITTI and NuScenes, to achieve state-of-the-art results with considerably less human supervision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a23153bf-93aa-4d15-9b39-b8e5f3e6fee8Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Structured Bird's-Eye-View Traffic Scene Understanding from Onboard ImagesYigit Baran Can, Alexander Liniger, Danda Pani Paudel, Luc Van GoolICCV 2021 · 147 citations
- Object-Driven Multi-Layer Scene Decomposition From a Single ImageHelisa Dhamo, Nassir Navab, Federico TombariICCV 2019 · 40 citations
- Self-Supervised Scene De-OcclusionXiaohang Zhan, Xingang Pan, Bo Dai, Ziwei Liu et al.CVPR 2020
- Projecting Your View Attentively: Monocular Road Scene Layout Estimation via Cross-View TransformationWeixiang Yang, Qi Li, Wenxi Liu, Yuanlong Yu et al.CVPR 2021
- Predicting Semantic Map Representations From Images Using Pyramid Occupancy NetworksThomas Roddick, Roberto CipollaCVPR 2020
Related papers
- MapPrior: Bird's-Eye View Map Layout Estimation with Generative ModelsXiyue Zhu, Vlas Zyrianov, Zhijian Liu, Shenlong WangICCV 2023 · 20 citations
- BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective SupervisionChenyu Yang, Yuntao Chen, Hao Tian, Chenxin Tao et al.CVPR 2023
- SkyEye: Self-Supervised Bird's-Eye-View Semantic Mapping Using Monocular Frontal View ImagesNikhil Gosala, Kürsat Petek, Paulo L. J. Drews-Jr, Wolfram Burgard et al.CVPR 2023
- OccluBEV: Occlusion Aware Spatiotemporal Modeling for Multi-view 3D Object DetectionZiteng Wen, Hai Xu, Chenyu Liu, Tao Guo et al.ACM MM 2023 · 5 citations
- Parametric Depth Based Feature Representation Learning for Object Detection and Segmentation in Bird's-Eye ViewJiayu Yang, Enze Xie, Miaomiao Liu, José M. ÁlvarezICCV 2023 · 9 citations
