Unleashing Semantic and Geometric Priors for 3D Scene Completion
Shiyuan Chen, Wei Sui, Bohao Zhang, Zeyd Boukhers, John See, Cong Yang
摘要
Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving and robotic navigation. However, existing methods rely on a coupled encoder to deliver both semantic and geometric priors, which forces the model to make a trade-off between conflicting demands and limits its overall performance. To tackle these challenges, we propose Foundation-SSC, a novel framework that performs dual decoupling at both the source and pathway levels. At the source level, we introduce a foundation encoder that provides rich semantic feature priors for the semantic branch and high-fidelity stereo cost volumes for the geometric branch. At the pathway level, these priors are refined through specialised, decoupled pathways, yielding superior semantic context and depth distributions. Our dual-decoupling design produces disentangled and refined inputs, which are then utilised by a hybrid view transformation to generate complementary 3D features. Additionally, we introduce a novel Axis-Aware Fusion (AAF) module that addresses the often-overlooked challenge of fusing these features by anisotropically merging them into a unified representation. Extensive experiments demonstrate the advantages of FoundationSSC, achieving simultaneous improvements in both semantic and geometric metrics, surpassing prior bests by +0.23 mIoU and +2.03 IoU on SemanticKITTI. Additionally, we achieve state-of-the-art performance on SSCBench-KITTI-360, with 21.78 mIoU and 48.61 IoU. Code - https://github.com/D-Robotics-AI-Lab/FoundationSSC * Equal contribution. Work done during Shiyuan Chen's internship at D-Robotics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
相关 Paper
- HD²-SSC: High-Dimension High-Density Semantic Scene Completion for Autonomous DrivingZhiwen Yang, Yuxin PengAAAI 2026
- Disentangling Instance and Scene Contexts for 3D Semantic Scene CompletionEnyu Liu, En Yu, Sijia Chen, Wenbing TaoICCV 2025 · 被引用 2 次
- Bi-SSC: Geometric-Semantic Bidirectional Fusion for Camera-Based 3D Semantic Scene CompletionYujie Xue, Ruihui Li, Fan Wu, Zhuo Tang 等CVPR 2024 · 被引用 8 次
- Learning Temporal 3D Semantic Scene Completion via Optical Flow GuidanceMeng Wang, Fan Wu, Ruihui Li, Yunchuan Qin 等NeurIPS 2025 · 被引用 4 次
- Symphonize 3D Semantic Scene Completion with Contextual Instance QueriesHaoyi Jiang, Tianheng Cheng, Naiyu Gao, Haoyang Zhang 等CVPR 2024
