FusionOcc: Multi-Modal Fusion for 3D Occupancy Prediction
Shuo Zhang, Yupeng Zhai, Jilin Mei, Yu Hu
摘要
3D occupancy prediction (OCC) aims to estimate and predict the semantic occupancy state of the surrounding environment, which is crucial for scene understanding and reconstruction in the real world. However, existing methods for 3D OCC mainly rely on surround-view camera images, whose performance is still insufficient in some challenging scenarios, such as low-light conditions. To this end, we propose a new multi-modal fusion network for 3D occupancy prediction by fusing features of LiDAR point clouds and surround-view images, called FusionOcc. Our model fuses features of these two modals in 2D and 3D space, respectively. By integrating the depth information from point clouds, a cross-modal fusion module is designed to predict a 2D dense depth map, enabling an accurate depth estimation and a better transition of 2D image features into 3D space. In addition, features of voxelized point clouds are aligned and merged with image features converted by a view-transformer in 3D space. Experiments show that FusionOcc establishes the new state of the art on Occ3D-nuScenes dataset, achieving a mIoU score of 35.94% (without visibility mask) and 56.62% (with visibility mask), showing an average improvement of 3.42% compared to the best previous method. Our work provides a new baseline for further research in multi-modal fusion for 3D occupancy prediction. Codes will be made publicly at https://github.com/ShuoZhang-code/FusionOcc.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action ModelXingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma 等AAAI 2026 · 被引用 119 次
- SAM4D: Segment Anything in Camera and LiDAR StreamsJianyun Xu, Song Wang, Ziqian Ni, Chunyong Hu 等ICCV 2025 · 被引用 2 次
- ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow PredictionsDubing Chen, Jin Fang, Wencheng Han, Xinjing Cheng 等ICCV 2025 · 被引用 2 次
- GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion PerceptionXiao Zhao, Chang Liu, Mingxu Zhu, Zheyuan Zhang 等ICLR 2026 · 被引用 1 次
- OccMamba: Semantic Occupancy Prediction with State Space ModelsHeng Li, Yuenan Hou, Xiaohan Xing, Yuexin Ma 等CVPR 2025
相关 Paper
- SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous DrivingYi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu 等ICCV 2023 · 被引用 380 次
- SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy PredictionZaipeng Duan, Chenxu Dang, Xuzhong Hu, Pei An 等CVPR 2025
- RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy PredictionBaojie Fan, Xiaotian Li, Yuhan Zhou, Yuyu Jiang 等ICCV 2025 · 被引用 1 次
- GAFusion: Adaptive Fusing LiDAR and Camera with Multiple Guidance for 3D Object DetectionXiaotian Li, Baojie Fan, Jiandong Tian, Huijie FanCVPR 2024
- POP-3D: Open-Vocabulary 3D Occupancy Prediction from ImagesAntonín Vobecký, Oriane Siméoni, David Hurych, Spyridon Gidaris 等NeurIPS 2023 · 被引用 67 次
