VoxDet: Rethinking 3D Semantic Scene Completion as Dense Object Detection
Wuyang Li, Zhu Yu, Alexandre Alahi
摘要
Semantic Scene Completion (SSC) aims to reconstruct the 3D geometry and semantics of the surrounding environment. With dense voxel labels, prior works typically formulate SSC as a dense segmentation task, independently classifying each voxel. However, this paradigm neglects critical instance-centric discriminability, leading to instance-level incompleteness and adjacent ambiguities. To address this, we highlight a "free lunch" of SSC labels: the voxel-level class label has implicitly told the instance-level insight, which is ever-overlooked by the community. Motivated by this observation, we first introduce a training-free Voxel-to-Instance (VoxNT) trick: a simple yet effective method that freely converts voxel-level class labels into instance-level offset labels. Building on this, we further propose VoxDet, an instance-centric framework that reformulates the voxel-level SSC as dense object detection by decoupling it into two sub-tasks: offset regression and semantic prediction. Specifically, based on the lifted 3D volume, VoxDet first uses (a) Spatially-decoupled Voxel Encoder to generate disentangled feature volumes for the two sub-tasks, which learn task-specific spatial deformation in the densely projected tri-perceptive space. Then, we deploy (b) Task-decoupled Dense Predictor to address SSC via dense detection. Here, we first regress a 4D offset field to estimate distances (6 directions) between voxels and the corresponding object boundaries in the voxel space. The regressed offsets are then used to guide the instance-level aggregation in the classification branch, achieving instance-aware scene completion. VoxDet can be deployed on both camera and LiDAR input and jointly achieves state-of-the-art results on both benchmarks, which gives 63.0 IoU on the SemanticKITTI test set, ranking 1 st on the online leaderboard.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper43
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei 等CVPR 2024 · 被引用 3,046 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
- Unifying Voxel-based Representation with Transformer for 3D Object DetectionYanwei Li, Yilun Chen, Xiaojuan Qi, Zeming Li 等NeurIPS 2022 · 被引用 401 次
相关 Paper
- Disentangling Instance and Scene Contexts for 3D Semantic Scene CompletionEnyu Liu, En Yu, Sijia Chen, Wenbing TaoICCV 2025 · 被引用 2 次
- VoxDet: Voxel Learning for Novel Instance DetectionBowen Li, Jiashun Wang, Yaoyu Hu, Chen Wang 等NeurIPS 2023 · 被引用 14 次
- 3D Instance Segmentation via Multi-Task Metric LearningJean Lahoud, Bernard Ghanem, Martin R. Oswald, Marc PollefeysICCV 2019 · 被引用 189 次
- H2GFormer: Horizontal-to-Global Voxel Transformer for 3D Semantic Scene CompletionYu Wang, Chao TongAAAI 2024 · 被引用 33 次
- HD²-SSC: High-Dimension High-Density Semantic Scene Completion for Autonomous DrivingZhiwen Yang, Yuxin PengAAAI 2026
