RadOcc: Learning Cross-Modality Occupancy Knowledge through Rendering Assisted Distillation
Haiming Zhang, Xu Yan, Dongfeng Bai, Jiantao Gao, Pan Wang, Bingbing Liu, Shuguang Cui, Zhen Li
Abstract
3D occupancy prediction is an emerging task that aims to estimate the occupancy states and semantics of 3D scenes using multi-view images. However, image-based scene perception encounters significant challenges in achieving accurate prediction due to the absence of geometric priors. In this paper, we address this issue by exploring cross-modal knowledge distillation in this task, i.e., we leverage a stronger multi-modal model to guide the visual model during training. In practice, we observe that directly applying features or logits alignment, proposed and widely used in bird's-eye-view (BEV) perception, does not yield satisfactory results. To overcome this problem, we introduce RadOcc, a Rendering assisted distillation paradigm for 3D Occupancy prediction. By employing differentiable volume rendering, we generate depth and semantic maps in perspective views and propose two novel consistency criteria between the rendered outputs of teacher and student models. Specifically, the depth consistency loss aligns the termination distributions of the rendered rays, while the semantic consistency loss mimics the intra-segment similarity guided by vision foundation models (VLMs). Experimental results on the nuScenes dataset demonstrate the effectiveness of our proposed method in improving various 3D occupancy prediction approaches, e.g., our proposed methodology enhances our baseline by 2.2% in the metric of mIoU and achieves 50% in Occ3D benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 598a50fb-284d-4a6f-aea4-98cf7a0c34acCited by top-tier papers11
- OctreeOcc: Efficient and Multi-Granularity Occupancy Prediction Using Octree QueriesYuhang Lu, Xinge Zhu, Tai Wang, Yuexin MaNeurIPS 2024 · 70 citations
- Learning 2D Invariant Affordance Knowledge for 3D Affordance GroundingXianqiang Gao, Pingrui Zhang, Delin Qu, Dong Wang et al.AAAI 2025 · 20 citations
- ProtoOcc: Accurate, Efficient 3D Occupancy Prediction Using Dual Branch Encoder-Prototype Query DecoderJungho Kim, Changwon Kang, Dongyoung Lee, Sehwan Choi et al.AAAI 2025 · 16 citations
- AMap: Distilling Future Priors for Ahead-Aware Online HD Map ConstructionRuikai Li, Xinrun Li, Mengwei Xie, Hao Shan et al.CVPR 2026 · 8 citations
- SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous DrivingHaiming Zhang, Yiyao Zhu, Wending Zhou, Xu Yan et al.NeurIPS 2025 · 5 citations
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous DrivingYi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu et al.ICCV 2023 · 380 citations
- Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene CompletionXu Yan, Jiantao Gao, Jie Li, Ruimao Zhang et al.AAAI 2021 · 365 citations
Related papers
- RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy PredictionBaojie Fan, Xiaotian Li, Yuhan Zhou, Yuyu Jiang et al.ICCV 2025 · 1 citation
- OccluBEV: Occlusion Aware Spatiotemporal Modeling for Multi-view 3D Object DetectionZiteng Wen, Hai Xu, Chenyu Liu, Tao Guo et al.ACM MM 2023 · 5 citations
- SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy PredictionZaipeng Duan, Chenxu Dang, Xuzhong Hu, Pei An et al.CVPR 2025
- X3KD: Knowledge Distillation Across Modalities, Tasks and Stages for Multi-Camera 3D Object DetectionMarvin Klingner, Shubhankar Borse, Varun Ravi Kumar, Behnaz Rezaei et al.CVPR 2023
- ViPOcc: Leveraging Visual Priors from Vision Foundation Models for Single-View 3D Occupancy PredictionYi Feng, Yu Han, Xijing Zhang, Tanghui Li et al.AAAI 2025 · 8 citations
