BAEFormer: Bi-Directional and Early Interaction Transformers for Bird's Eye View Semantic Segmentation
Cong Pan, Yonghao He, Junran Peng, Qian Zhang, Wei Sui, Zhaoxiang Zhang
Abstract
Bird's Eye View (BEV) semantic segmentation is a critical task in autonomous driving. However, existing Transformer-based methods confront difficulties in transforming Perspective View (PV) to BEV due to their unidirectional and posterior interaction mechanisms. To address this issue, we propose a novel Bi-directional and Early Interaction Transformers framework named BAEFormer, consisting of (i) an early-interaction PV-BEV pipeline and (ii) a bi-directional cross-attention mechanism. Moreover, we find that the image feature maps' resolution in the crossattention module has a limited effect on the final performance. Under this critical observation, we propose to enlarge the size of input images and downsample the multiview image features for cross-interaction, further improving the accuracy while keeping the amount of computation controllable. Our proposed method for BEV semantic segmentation achieves state-of-the-art performance in real-time inference speed on the nuScenes dataset, i.e., 38.9 mIoU at 45 FPS on a single A100 GPU.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62a43d98-55e6-4e6b-a74c-220180cbd146Cited by top-tier papers7
- HTNav: A Hybrid Navigation Framework with Tiered Structure for Urban Aerial Vision-and-Language NavigationChengjie Fan, Cong Pan, Zijian Liu, Ningzhong Liu et al.CVPR 2026 · 4 citations
- History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language NavigationXichen Ding, Jianzhe Gao, Cong Pan, Wenguan Wang et al.AAAI 2026 · 3 citations
- ProOOD: Prototype-Guided Out-of-Distribution 3D Occupancy PredictionYuheng Zhang, Mengfei Duan, Kunyu Peng, Yuhang Wang et al.CVPR 2026
- InteractionMap: Improving Online Vectorized HDMap Construction with InteractionKuang Wu, Chuan Yang, Zhanbin LiCVPR 2025
- Seeing in Double: Dual-Granularity BEV Segmentation via Mamba-Driven Alignment and Polar-Decoupled ExpertsJiaxin Cai, Rui Lin, Jingze Su, Qi Li et al.AAAI 2026
Builds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li et al.ICCV 2023 · 513 citations
- FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular CamerasAnthony Hu, Zak Murez, Nikhil Mohan, Sofía Dudas et al.ICCV 2021 · 329 citations
- Cross-view Transformers for real-time Map-view Semantic SegmentationBrady Zhou, Philipp KrähenbühlCVPR 2022 · 279 citations
- NEAT: Neural Attention Fields for End-to-End Autonomous DrivingKashyap Chitta, Aditya Prakash, Andreas GeigerICCV 2021 · 274 citations
Related papers
- CycleBEV: Regularizing View Transformation Networks via View Cycle Consistency for Bird’s-Eye-View Semantic SegmentationJeongbin Hong, Dooseop Choi, Taeg-Hyun An, KYOUNG AN AN et al.CVPR 2026 · 1 citation
- PolarFormer: Multi-Camera 3D Object Detection with Polar TransformerYanqin Jiang, Li Zhang, Zhenwei Miao, Xiatian Zhu et al.AAAI 2023 · 240 citations
- Tri-Perspective View for Vision-Based 3D Semantic Occupancy PredictionYuanhui Huang, Wenzhao Zheng, Yunpeng Zhang, Jie Zhou et al.CVPR 2023
- TBP-Former: Learning Temporal Bird's-Eye-View Pyramid for Joint Perception and Prediction in Vision-Centric Autonomous DrivingShaoheng Fang, Zi Wang, Yiqi Zhong, Junhao Ge et al.CVPR 2023
- OccluBEV: Occlusion Aware Spatiotemporal Modeling for Multi-view 3D Object DetectionZiteng Wen, Hai Xu, Chenyu Liu, Tao Guo et al.ACM MM 2023 · 5 citations
