MatrixVT: Efficient Multi-Camera to BEV Transformation for 3D Perception
Hongyu Zhou, Zheng Ge, Zeming Li, Xiangyu Zhang
Abstract
This paper proposes an efficient multi-camera to Bird's-Eye-View (BEV) view transformation method for 3D perception, dubbed MatrixVT. Existing view transformers either suffer from poor transformation efficiency or rely on device-specific operators, hindering the broad application of BEV models. In contrast, our method generates BEV features efficiently with only convolutions and matrix multiplications (MatMul). Specifically, we propose describing the BEV feature as the MatMul of image feature and a sparse Feature Transporting Matrix (FTM). A Prime Extraction module is then introduced to compress the dimension of image features and reduce FTM's sparsity. Moreover, we propose the Ring & Ray Decomposition to replace the FTM with two matrices and reformulate our pipeline to reduce calculation further. Compared to existing methods, MatrixVT enjoys a faster speed and less memory footprint while remaining deploy-friendly. Extensive experiments on the nuScenes benchmark demonstrate that our method is highly efficient but obtains results on par with the SOTA method in object detection and map segmentation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6af870c8-ace2-4bcf-bfb9-721e31b1201aCited by top-tier papers5
- OctOcc: High-Resolution 3D Occupancy Prediction with OctreeWenzhe Ouyang, Xiaolin Song, Bailan Feng, Zenglin XuAAAI 2024 · 12 citations
- RayFormer: Improving Query-Based Multi-Camera 3D Object Detection via Ray-Centric StrategiesXiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan et al.ACM MM 2024 · 7 citations
- FastRSR: Efficient and Accurate Road Surface Reconstruction in Bird's Eye ViewYuting Zhao, Yuheng Ji, Xiaoshuai Hao, Shuxiao LiACM MM 2025 · 2 citations
- PointBeV: A Sparse Approach to BeV PredictionsLoïck Chambon, Éloi Zablocki, Mickaël Chen, Florent Bartoccioni et al.CVPR 2024
- Spe-BEVHead: Rethinking the Detection Head Design for Bird's-Eye-View Object DetectionJunshu Zhang, Sicheng Zhao, Xin Zhao, Fan Yang et al.CVPR 2026
Builds on10
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang et al.CVPR 2022 · 794 citations
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li et al.ICCV 2021 · 404 citations
- FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular CamerasAnthony Hu, Zak Murez, Nikhil Mohan, Sofía Dudas et al.ICCV 2021 · 329 citations
Related papers
- EVT: Efficient View Transformation for Multi-Modal 3D Object DetectionYongjin Lee, Hyeon Mun Jeong, Yurim Jeon, Sanghyun KimICCV 2025 · 5 citations
- FB-BEV: BEV Representation from Forward-Backward View TransformationsZhiqi Li, Zhiding Yu, Wenhai Wang, Anima Anandkumar et al.ICCV 2023 · 144 citations
- MaskBEV: Towards A Unified Framework for BEV Detection and Map SegmentationXiao Zhao, Xukun Zhang, Dingkang Yang, Mingyang Sun et al.ACM MM 2024 · 7 citations
- OccluBEV: Occlusion Aware Spatiotemporal Modeling for Multi-view 3D Object DetectionZiteng Wen, Hai Xu, Chenyu Liu, Tao Guo et al.ACM MM 2023 · 5 citations
- BEVNeXt: Reviving Dense BEV Frameworks for 3D Object DetectionZhenxin Li, Shiyi Lan, José M. Álvarez, Zuxuan WuCVPR 2024
