Projecting Your View Attentively: Monocular Road Scene Layout Estimation via Cross-View Transformation
Weixiang Yang, Qi Li, Wenxi Liu, Yuanlong Yu, Yuexin Ma, Shengfeng He, Jia Pan
Abstract
HD map reconstruction is crucial for autonomous driving. LiDAR-based methods are limited due to the deployed expensive sensors and time-consuming computation. Camera-based methods usually need to separately perform road segmentation and view transformation, which often causes distortion and the absence of content. To push the limits of the technology, we present a novel framework that enables reconstructing a local map formed by road layout and vehicle occupancy in the bird's-eye view given a front-view monocular image only. In particular, we propose a cross-view transformation module, which takes the constraint of cycle consistency between views into account and makes full use of their correlation to strengthen the view transformation and scene understanding. Considering the relationship between vehicles and roads, we also design a context-aware discriminator to further refine the results. Experiments on public benchmarks show that our method achieves the state-of-the-art performance in the tasks of road layout estimation and vehicle occupancy estimation. Especially for the latter task, our model outperforms all competitors by a large margin. Furthermore, our model runs at 35 FPS on a single GPU, which is efficient and applicable for real-time panorama HD map reconstruction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2074b17-415d-428e-926a-6cd083b2c382Cited by top-tier papers16
- VectorMapNet: End-to-end Vectorized HD Map LearningYicheng Liu, Tianyuan Yuan, Yue Wang, Yilun Wang et al.ICML 2023 · 332 citations
- PolarFormer: Multi-Camera 3D Object Detection with Polar TransformerYanqin Jiang, Li Zhang, Zhenwei Miao, Xiatian Zhu et al.AAAI 2023 · 240 citations
- PivotNet: Vectorized Pivot Learning for End-to-end HD Map ConstructionWenjie Ding, Limeng Qiao, Xi Qiu, Chi ZhangICCV 2023 · 119 citations
- DiffBEV: Conditional Diffusion Model for Bird's Eye View PerceptionJiayu Zou, Kun Tian, Zheng Zhu, Yun Ye et al.AAAI 2024 · 42 citations
- JOTR: 3D Joint Contrastive Learning with Transformers for Occluded Human Mesh RecoveryJiahao Li, Zongxin Yang, Xiaohan Wang, Jianxin Ma et al.ICCV 2023 · 22 citations
Builds on10
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Learning Lightweight Lane Detection CNNs by Self Attention DistillationYuenan Hou, Zheng Ma, Chunxiao Liu, Chen Change LoyICCV 2019 · 666 citations
- Disentangling Monocular 3D Object DetectionAndrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Manuel Lopez-Antequera et al.ICCV 2019 · 504 citations
- Advisable Learning for Self-Driving Vehicles by Internalizing Observation-to-Action RulesJinkyu Kim, Suhong Moon, Anna Rohrbach, Trevor Darrell et al.CVPR 2020
Related papers
- SafeMap: Robust HD Map Construction from Incomplete ObservationsXiaoshuai Hao, Lingdong Kong, Rong Yin, Pengwei Wang et al.ICML 2025
- Predicting Semantic Map Representations From Images Using Pyramid Occupancy NetworksThomas Roddick, Roberto CipollaCVPR 2020
- SkyEye: Self-Supervised Bird's-Eye-View Semantic Mapping Using Monocular Frontal View ImagesNikhil Gosala, Kürsat Petek, Paulo L. J. Drews-Jr, Wolfram Burgard et al.CVPR 2023
- COTR: Compact Occupancy TRansformer for Vision-Based 3D Occupancy PredictionQihang Ma, Xin Tan, Yanyun Qu, Lizhuang Ma et al.CVPR 2024 · 30 citations
- DVGT: Driving Visual Geometry TransformerSicheng Zuo, Zixun Xie, Wenzhao Zheng, Shaoqing Xu et al.CVPR 2026 · 23 citations
