ProtoOcc: Accurate, Efficient 3D Occupancy Prediction Using Dual Branch Encoder-Prototype Query Decoder
Jungho Kim, Changwon Kang, Dongyoung Lee, Sehwan Choi, Jun Won Choi
摘要
In this paper, we introduce ProtoOcc, a novel 3D occupancy prediction model designed to predict the occupancy states and semantic classes of 3D voxels via a deep semantic understanding of scenes. ProtoOcc consists of two main components: the Dual Branch Encoder (DBE) and the Prototype Query Decoder (PQD). The DBE produces a new 3D voxel representation by combining 3D voxel and BEV representations across multiple scales using a dual branch structure. This design combines the BEV representation, which offers a large receptive field, with the voxel representation, known for its higher spatial resolution, thereby improving both performance and computational efficiency. The PQD employs two types of prototype-based queries to expedite the Transformer decoding process. Scene-Adaptive Prototypes are generated from the 3D voxel features of the input sample, while Scene-Agnostic Prototypes are updated during training using an Exponential Moving Average of the Scene-Adaptive Prototypes. Using these prototype-based queries for decoding, we can directly predict 3D occupancy in a single step, eliminating the need for iterative Transformer decoding. Additionally, we propose Robust Prototype Learning, which introduces noise into the prototype generation process and trains the model to denoise during the training phase. This approach enhances the robustness of ProtoOcc against degraded prototype feature quality. ProtoOcc achieves state-of-the-art performance with 45.02% mIoU on the Occ3D-nuScenes benchmark. For the single-frame method, it reaches 39.56% mIoU with 12.83 FPS on an NVIDIA RTX 3090. Our code can be found at https://github.com/SPA-junghokim/ProtoOcc .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- See through the Dark: Learning Illumination-affined Representations for Nighttime Occupancy PredictionYuan Wu, Zhiqiang Yan, Yigong Zhang, Xiang Li 等NeurIPS 2025 · 被引用 7 次
- SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic QueriesChenxu Dang, Haiyan Liu, Jason Bao, Pei An 等AAAI 2026 · 被引用 6 次
- Semantic Causality-Aware Vision-Based 3D Occupancy PredictionDubing Chen, Huan Zheng, Yucheng Zhou, Xianfei Li 等ICCV 2025 · 被引用 1 次
- ProOOD: Prototype-Guided Out-of-Distribution 3D Occupancy PredictionYuheng Zhang, Mengfei Duan, Kunyu Peng, Yuhang Wang 等CVPR 2026
- MAESTRO: Task-Relevant Optimization Via Adaptive Feature Enhancement and Suppression for Multi-Task 3D PerceptionChangwon Kang, Jisong Kim, Hongjae Shin, Junseo Park 等ICCV 2025
它引用的顶会 Paper23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
- SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous DrivingYi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu 等ICCV 2023 · 被引用 380 次
相关 Paper
- 3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View TransformationGyeongrok Oh, Sungjune Kim, Heeju Ko, Hyung-gun Chi 等CVPR 2025
- OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy PredictionYunpeng Zhang, Zheng Zhu, Dalong DuICCV 2023 · 被引用 354 次
- RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy PredictionBaojie Fan, Xiaotian Li, Yuhan Zhou, Yuyu Jiang 等ICCV 2025 · 被引用 1 次
- OPUS: Occupancy Prediction Using a Sparse SetJiabao Wang, Zhaojiang Liu, Qiang Meng, Liujiang Yan 等NeurIPS 2024 · 被引用 67 次
- Achieving Speed-Accuracy Balance in Vision-based 3D Occupancy Prediction via Geometric-Semantic DisentanglementYulin He, Wei Chen, Siqi Wang, Tianci Xun 等AAAI 2025 · 被引用 4 次
