ProtoOcc: Accurate, Efficient 3D Occupancy Prediction Using Dual Branch Encoder-Prototype Query Decoder
Jungho Kim, Changwon Kang, Dongyoung Lee, Sehwan Choi, Jun Won Choi
Abstract
In this paper, we introduce ProtoOcc, a novel 3D occupancy prediction model designed to predict the occupancy states and semantic classes of 3D voxels via a deep semantic understanding of scenes. ProtoOcc consists of two main components: the Dual Branch Encoder (DBE) and the Prototype Query Decoder (PQD). The DBE produces a new 3D voxel representation by combining 3D voxel and BEV representations across multiple scales using a dual branch structure. This design combines the BEV representation, which offers a large receptive field, with the voxel representation, known for its higher spatial resolution, thereby improving both performance and computational efficiency. The PQD employs two types of prototype-based queries to expedite the Transformer decoding process. Scene-Adaptive Prototypes are generated from the 3D voxel features of the input sample, while Scene-Agnostic Prototypes are updated during training using an Exponential Moving Average of the Scene-Adaptive Prototypes. Using these prototype-based queries for decoding, we can directly predict 3D occupancy in a single step, eliminating the need for iterative Transformer decoding. Additionally, we propose Robust Prototype Learning, which introduces noise into the prototype generation process and trains the model to denoise during the training phase. This approach enhances the robustness of ProtoOcc against degraded prototype feature quality. ProtoOcc achieves state-of-the-art performance with 45.02% mIoU on the Occ3D-nuScenes benchmark. For the single-frame method, it reaches 39.56% mIoU with 12.83 FPS on an NVIDIA RTX 3090. Our code can be found at https://github.com/SPA-junghokim/ProtoOcc .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3bf0d48f-fb5b-43fd-9acd-4febb23f8064Cited by top-tier papers6
- See through the Dark: Learning Illumination-affined Representations for Nighttime Occupancy PredictionYuan Wu, Zhiqiang Yan, Yigong Zhang, Xiang Li et al.NeurIPS 2025 · 7 citations
- SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic QueriesChenxu Dang, Haiyan Liu, Jason Bao, Pei An et al.AAAI 2026 · 6 citations
- Semantic Causality-Aware Vision-Based 3D Occupancy PredictionDubing Chen, Huan Zheng, Yucheng Zhou, Xianfei Li et al.ICCV 2025 · 1 citation
- ProOOD: Prototype-Guided Out-of-Distribution 3D Occupancy PredictionYuheng Zhang, Mengfei Duan, Kunyu Peng, Yuhang Wang et al.CVPR 2026
- MAESTRO: Task-Relevant Optimization Via Adaptive Feature Enhancement and Suppression for Multi-Task 3D PerceptionChangwon Kang, Jisong Kim, Hongjae Shin, Junseo Park et al.ICCV 2025
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous DrivingYi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu et al.ICCV 2023 · 380 citations
Related papers
- 3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View TransformationGyeongrok Oh, Sungjune Kim, Heeju Ko, Hyung-gun Chi et al.CVPR 2025
- OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy PredictionYunpeng Zhang, Zheng Zhu, Dalong DuICCV 2023 · 354 citations
- RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy PredictionBaojie Fan, Xiaotian Li, Yuhan Zhou, Yuyu Jiang et al.ICCV 2025 · 1 citation
- OPUS: Occupancy Prediction Using a Sparse SetJiabao Wang, Zhaojiang Liu, Qiang Meng, Liujiang Yan et al.NeurIPS 2024 · 67 citations
- Achieving Speed-Accuracy Balance in Vision-based 3D Occupancy Prediction via Geometric-Semantic DisentanglementYulin He, Wei Chen, Siqi Wang, Tianci Xun et al.AAAI 2025 · 4 citations
