Improving Bird's Eye View Semantic Segmentation by Task Decomposition
Tianhao Zhao, Yongcan Chen, Yu Wu, Tianyang Liu, Bo Du, Peilun Xiao, Shi Qiu, Hongda Yang, Guozhen Li, Yi Yang, Yutian Lin
摘要
Semantic segmentation in bird's eye view (BEV) plays a crucial role in autonomous driving. Previous methods usually follow an end-to-end pipeline, directly predicting the BEV segmentation map from monocular RGB inputs. However, the challenge arises when the RGB inputs and BEV targets from distinct perspectives, making the direct point-to-point predicting hard to optimize. In this paper, we decompose the original BEV segmentation task into two stages, namely BEV map reconstruction and RGB-BEV feature alignment. In the first stage, we train a BEV autoencoder to reconstruct the BEV segmentation maps given corrupted noisy latent representation, which urges the decoder to learn fundamental knowledge of typical BEV patterns. The second stage involves mapping RGB input images into the BEV latent space of the first stage, directly optimizing the correlations between the two views at the feature level. Our approach simplifies the complexity of combining perception and generation into distinct steps, equipping the model to handle intricate and challenging scenes effectively. Besides, we propose to transform the BEV segmentation map from the Cartesian to the polar coordinate system to establish the column-wise correspondence between RGB images and BEV maps. Moreover, our method requires neither multi-scale features nor camera intrinsic parameters for depth estimation and saves computational overhead. Extensive experiments on nuScenes and Argoverse show the effectiveness and efficiency of our method. Code is available at https://github.com/happytianhao/TaDe . * Equal contribution. † Corresponding author. (a) Perspective View RGB Image (b) End-to-end (c) TaDe (Ours)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- RayFormer: Improving Query-Based Multi-Camera 3D Object Detection via Ray-Centric StrategiesXiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan 等ACM MM 2024 · 被引用 7 次
- RiskProp: Collision-Anchored Self-Supervised Risk Propagation For Early Accident AnticipationYiyang Zou, Tianhao Zhao, Peilun Xiao, Hongyu Jin 等CVPR 2026 · 被引用 4 次
- VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector QuantizationYiwei Zhang, Jin Gao, Fudong Ge, Guan Luo 等NeurIPS 2024 · 被引用 3 次
- CycleBEV: Regularizing View Transformation Networks via View Cycle Consistency for Bird’s-Eye-View Semantic SegmentationJeongbin Hong, Dooseop Choi, Taeg-Hyun An, KYOUNG AN AN 等CVPR 2026 · 被引用 1 次
- Monocular Vehicle Pose and Shape Reconstruction via Dynamic Context Adaptation and Progressive Geometry RefinementWei Li, Long Ji, Ying Wang, Xiao Wu 等AAAI 2026
它引用的顶会 Paper14
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real 等NeurIPS 2023 · 被引用 734 次
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li 等ICCV 2023 · 被引用 513 次
- FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular CamerasAnthony Hu, Zak Murez, Nikhil Mohan, Sofía Dudas 等ICCV 2021 · 被引用 329 次
- Cross-view Transformers for real-time Map-view Semantic SegmentationBrady Zhou, Philipp KrähenbühlCVPR 2022 · 被引用 279 次
相关 Paper
- CalibRBEV: Multi-Camera Calibration via Reversed Bird's-eye-view Representations for Autonomous DrivingWenlong Liao, Sunyuan Qiang, Xianfei Li, Xiaolei Chen 等ACM MM 2024 · 被引用 2 次
- BEV-CAR: Enhancing Monocular Bird's Eye View Segmentation with Context-Aware RasterizationYixin Xiong, Ke Wang, Tongtong Cheng, Chunhui Liu 等CVPR 2026
- ARINBEV: Bird's-Eye View Layout Estimation with Conditional Autoregressive ModelJiyong Kwag, Charles K. Toth, Alper YilmazICLR 2026
- SkyEye: Self-Supervised Bird's-Eye-View Semantic Mapping Using Monocular Frontal View ImagesNikhil Gosala, Kürsat Petek, Paulo L. J. Drews-Jr, Wolfram Burgard 等CVPR 2023
- SA-BEV: Generating Semantic-Aware Bird's-Eye-View Feature for Multi-view 3D Object DetectionJinqing Zhang, Yanan Zhang, Qingjie Liu, Yunhong WangICCV 2023 · 被引用 41 次
