FB-BEV: BEV Representation from Forward-Backward View Transformations
Zhiqi Li, Zhiding Yu, Wenhai Wang, Anima Anandkumar, Tong Lu, José M. Álvarez
摘要
View Transformation Module (VTM), where transformations happen between multi-view image features and Bird-Eye-View (BEV) representation, is a crucial step in camera-based BEV perception systems. Currently, the two most prominent VTM paradigms are forward projection and backward projection. Forward projection, represented by Lift-Splat-Shoot, leads to sparsely projected BEV features without post-processing. Backward projection, with BEV-Former being an example, tends to generate false-positive BEV features from incorrect projections due to the lack of utilization on depth. To address the above limitations, we propose a novel forward-backward view transformation module. Our approach compensates for the deficiencies in both existing methods, allowing them to enhance each other to obtain higher quality BEV representations mutually. We instantiate the proposed module with FB-BEV, which achieves a new state-of-the-art result of 62.4% NDS on the nuScenes test set. Code and models are available at https://github.com/NVlabs/FB-BEV.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- Memory-and-Anticipation Transformer for Online Action UnderstandingJiahao Wang, Guo Chen, Yifei Huang, Limin Wang 等ICCV 2023 · 被引用 72 次
- OctreeOcc: Efficient and Multi-Granularity Occupancy Prediction Using Octree QueriesYuhang Lu, Xinge Zhu, Tai Wang, Yuexin MaNeurIPS 2024 · 被引用 70 次
- COTR: Compact Occupancy TRansformer for Vision-Based 3D Occupancy PredictionQihang Ma, Xin Tan, Yanyun Qu, Lizhuang Ma 等CVPR 2024 · 被引用 30 次
- GeoBEV: Learning Geometric BEV Representation for Multi-view 3D Object DetectionJinqing Zhang, Yanan Zhang, Yunlong Qi, Zehua Fu 等AAAI 2025 · 被引用 22 次
- Regulating Intermediate 3D Features for Vision-Centric Autonomous DrivingJunkai Xu, Liang Peng, Haoran Cheng, Linxuan Xia 等AAAI 2024 · 被引用 14 次
它引用的顶会 Paper16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang 等CVPR 2022 · 被引用 794 次
相关 Paper
- MatrixVT: Efficient Multi-Camera to BEV Transformation for 3D PerceptionHongyu Zhou, Zheng Ge, Zeming Li, Xiangyu ZhangICCV 2023 · 被引用 62 次
- EVT: Efficient View Transformation for Multi-Modal 3D Object DetectionYongjin Lee, Hyeon Mun Jeong, Yurim Jeon, Sanghyun KimICCV 2025 · 被引用 5 次
- RayFormer: Improving Query-Based Multi-Camera 3D Object Detection via Ray-Centric StrategiesXiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan 等ACM MM 2024 · 被引用 7 次
- CycleBEV: Regularizing View Transformation Networks via View Cycle Consistency for Bird’s-Eye-View Semantic SegmentationJeongbin Hong, Dooseop Choi, Taeg-Hyun An, KYOUNG AN AN 等CVPR 2026 · 被引用 1 次
- BAEFormer: Bi-Directional and Early Interaction Transformers for Bird's Eye View Semantic SegmentationCong Pan, Yonghao He, Junran Peng, Qian Zhang 等CVPR 2023
