VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene Completion
Meng Wang, Huilong Pi, Ruihui Li, Yunchuan Qin, Zhuo Tang, Kenli Li
摘要
Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving. However, images provide limited information making the model susceptible to geometric ambiguity caused by occlusion and perspective distortion. Existing methods often lack explicit semantic modeling between objects, limiting their perception of 3D semantic context. To address these challenges, we propose a novel method VLScene: Vision-Language Guidance Distillation for Camera-based 3D Semantic Scene Completion. The key insight is to use the vision-language model to introduce high-level semantic priors to provide the object spatial context required for 3D scene understanding. Specifically, we design a vision-language guidance distillation process to enhance image features, which can effectively capture semantic knowledge from the surrounding environment and improve spatial context reasoning. In addition, we introduce a geometric-semantic sparse awareness mechanism to propagate geometric structures in the neighborhood and enhance semantic information through contextual sparse interactions. Experimental results demonstrate that VLScene achieves rank-1st performance on challenging benchmarks—SemanticKITTI and SSCBench-KITTI-360, yielding remarkably mIoU scores of 17.52 and 19.10, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Learning Temporal 3D Semantic Scene Completion via Optical Flow GuidanceMeng Wang, Fan Wu, Ruihui Li, Yunchuan Qin 等NeurIPS 2025 · 被引用 4 次
- SDFormer: Vision-Based 3D Semantic Scene Completion via SAM-Assisted Dual-Channel Voxel TransformerYujie Xue, Huilong Pi, Jiapeng Zhang, Yunchuan Qin 等ICCV 2025 · 被引用 3 次
- VoxDet: Rethinking 3D Semantic Scene Completion as Dense Object DetectionWuyang Li, Zhu Yu, Alexandre AlahiNeurIPS 2025 · 被引用 3 次
- Unleashing Semantic and Geometric Priors for 3D Scene CompletionShiyuan Chen, Wei Sui, Bohao Zhang, Zeyd Boukhers 等AAAI 2026 · 被引用 1 次
- Learning Spatial-Temporal Consistency for 3D Semantic Scene CompletionYujie Xue, Meng Wang, Ruihui Li, Fan Wu 等CVPR 2026
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun 等ICLR 2022 · 被引用 885 次
相关 Paper
- HD²-SSC: High-Dimension High-Density Semantic Scene Completion for Autonomous DrivingZhiwen Yang, Yuxin PengAAAI 2026
- Bi-SSC: Geometric-Semantic Bidirectional Fusion for Camera-Based 3D Semantic Scene CompletionYujie Xue, Ruihui Li, Fan Wu, Zhuo Tang 等CVPR 2024 · 被引用 8 次
- Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning SynergyHaijier Chen, Bo Xu, Shoujian Zhang, Haoze Liu 等ICLR 2026 · 被引用 6 次
- Distilling Diffusion Models to Efficient 3D LiDAR Scene CompletionShengyuan Zhang, An Zhao, Ling Yang, Zejian Li 等ICCV 2025 · 被引用 1 次
- Towards 3D Object-Centric Feature Learning for Semantic Scene CompletionWeihua Wang, Yubo Cui, Xiangru Lin, Zhiheng Li 等AAAI 2026
