FusionOcc: Multi-Modal Fusion for 3D Occupancy Prediction
Shuo Zhang, Yupeng Zhai, Jilin Mei, Yu Hu
Abstract
3D occupancy prediction (OCC) aims to estimate and predict the semantic occupancy state of the surrounding environment, which is crucial for scene understanding and reconstruction in the real world. However, existing methods for 3D OCC mainly rely on surround-view camera images, whose performance is still insufficient in some challenging scenarios, such as low-light conditions. To this end, we propose a new multi-modal fusion network for 3D occupancy prediction by fusing features of LiDAR point clouds and surround-view images, called FusionOcc. Our model fuses features of these two modals in 2D and 3D space, respectively. By integrating the depth information from point clouds, a cross-modal fusion module is designed to predict a 2D dense depth map, enabling an accurate depth estimation and a better transition of 2D image features into 3D space. In addition, features of voxelized point clouds are aligned and merged with image features converted by a view-transformer in 3D space. Experiments show that FusionOcc establishes the new state of the art on Occ3D-nuScenes dataset, achieving a mIoU score of 35.94% (without visibility mask) and 56.62% (with visibility mask), showing an average improvement of 3.42% compared to the best previous method. Our work provides a new baseline for further research in multi-modal fusion for 3D occupancy prediction. Codes will be made publicly at https://github.com/ShuoZhang-code/FusionOcc.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3a7524ad-6c8e-4f7b-9159-80ff32b1aed4Cited by top-tier papers5
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action ModelXingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma et al.AAAI 2026 · 119 citations
- SAM4D: Segment Anything in Camera and LiDAR StreamsJianyun Xu, Song Wang, Ziqian Ni, Chunyong Hu et al.ICCV 2025 · 2 citations
- ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow PredictionsDubing Chen, Jin Fang, Wencheng Han, Xinjing Cheng et al.ICCV 2025 · 2 citations
- GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion PerceptionXiao Zhao, Chang Liu, Mingxu Zhu, Zheyuan Zhang et al.ICLR 2026 · 1 citation
- OccMamba: Semantic Occupancy Prediction with State Space ModelsHeng Li, Yuenan Hou, Xiaohan Xing, Yuexin Ma et al.CVPR 2025
Related papers
- SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous DrivingYi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu et al.ICCV 2023 · 380 citations
- SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy PredictionZaipeng Duan, Chenxu Dang, Xuzhong Hu, Pei An et al.CVPR 2025
- RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy PredictionBaojie Fan, Xiaotian Li, Yuhan Zhou, Yuyu Jiang et al.ICCV 2025 · 1 citation
- GAFusion: Adaptive Fusing LiDAR and Camera with Multiple Guidance for 3D Object DetectionXiaotian Li, Baojie Fan, Jiandong Tian, Huijie FanCVPR 2024
- POP-3D: Open-Vocabulary 3D Occupancy Prediction from ImagesAntonín Vobecký, Oriane Siméoni, David Hurych, Spyridon Gidaris et al.NeurIPS 2023 · 67 citations
