SPSC: Sparse and Scalable Multi-Modal 3D Occupancy Prediction for Autonomous Driving
Qingju Guo, Shuang Li, Binhui Xie, Jing Geng, Wei Li
摘要
3D semantic occupancy prediction offers a nuanced representation of the surrounding environment, which is crucial for ensuring the safety of autonomous driving. However, fine-grained scene representations inevitably result in cubic growth in data scale, which imposes substantial demands on model architecture and computational complexity, especially in high-resolution scenarios. Existing approaches for handling high-resolution scenes typically obtain fine-grained features by grid sampling on low-resolution feature map, resulting in limited sparsity and insufficient feature interaction. This paper presents a framework leveraging SParse representation and SCalable feature interaction to address the aforementioned challenges, called SPSC. Specifically, we maintain sparsity by progressively pruning unoccupied queries during the coarse-to-fine process, thereby reducing the scale of data that the model needs to handle. Subsequently, we introduce query serialization, which transforms queries into an ordered sequence while preserving their spatial structure. This enables fine-grained feature interaction while maintaining linear computational complexity and a larger receptive field. Without complex architectural designs, SPSC significantly outperforms SOTA approaches, enhances the mIoU by 12.0%, 11.0% and 4.8% on nuScenes-Occupancy dataset under the muli-modal, LiDAR and camera settings, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- Conditional Positional Encodings for Vision TransformersXiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang 等ICLR 2023 · 被引用 406 次
- OpenOccupancy: A Large Scale Benchmark for Surrounding Semantic Occupancy PerceptionXiaofeng Wang, Zheng Zhu, Wenbo Xu, Yunpeng Zhang 等ICCV 2023 · 被引用 270 次
- MonoScene: Monocular 3D Semantic Scene CompletionAnh-Quan Cao, Raoul de CharetteCVPR 2022 · 被引用 251 次
相关 Paper
- OctOcc: High-Resolution 3D Occupancy Prediction with OctreeWenzhe Ouyang, Xiaolin Song, Bailan Feng, Zenglin XuAAAI 2024 · 被引用 12 次
- CSV-Occ: Fusing Multi-frame Alignment for Occupancy Prediction with Temporal Cross State Space Model and Central Voting MechanismZiming Zhu, Yu Zhu, Jiahao Chen, Xiaofeng Ling 等ICML 2025
- ODG: Occupancy Prediction Using Dual GaussiansYunxiao Shi, Yinhao Zhu, Herbert Cai, Shizhong Han 等NeurIPS 2025 · 被引用 7 次
- S2GO: Streaming Sparse Gaussian OccupancyJinhyung Park, Chensheng Peng, Yihan Hu, Wenzhao Zheng 等ICLR 2026 · 被引用 2 次
- Occupancy Learning with Spatiotemporal MemoryZiyang Leng, Jiawei Yang, Wenlong Yi, Bolei ZhouICCV 2025 · 被引用 10 次
