One at a Time: Progressive Multi-Step Volumetric Probability Learning for Reliable 3D Scene Perception
Bohan Li, Yasheng Sun, Jingxin Dong, Zheng Zhu, Jinming Liu, Xin Jin, Wenjun Zeng
摘要
Numerous studies have investigated the pivotal role of reliable 3D volume representation in scene perception tasks, such as multi-view stereo (MVS) and semantic scene completion (SSC). They typically construct 3D probability volumes directly with geometric correspondence, attempting to fully address the scene perception tasks in a single forward pass. However, such a single-step solution makes it hard to learn accurate and convincing volumetric probability, especially in challenging regions like unexpected occlusions and complicated light reflections. Therefore, this paper proposes to decompose the complicated 3D volume representation learning into a sequence of generative steps to facilitate fine and reliable scene perception. Considering the recent advances achieved by strong generative diffusion models, we introduce a multi-step learning framework, dubbed as VPD, dedicated to progressively refining the Volumetric Probability in a Diffusion process. Specifically, we first build a coarse probability volume from input images with the off-the-shelf scene perception baselines, which is then conditioned as the basic geometry prior before being fed into a 3D diffusion UNet, to progressively achieve accurate probability distribution modeling. To handle the corner cases in challenging areas, a Confidence-Aware Contextual Collaboration (CACC) module is developed to correct the uncertain regions for reliable volumetric learning based on multi-scale contextual contents. Moreover, an Online Filtering (OF) strategy is designed to maintain representation consistency for stable diffusion sampling. Extensive experiments are conducted on scene perception tasks including multi-view stereo (MVS) and semantic scene completion (SSC), to validate the efficacy of our method in learning reliable volumetric representations. Notably, for the SSC task, our work stands out as the first to surpass LiDAR-based methods on the SemanticKITTI dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Hybrid-Grained Feature Aggregation with Coarse-to-Fine Language Guidance for Self-Supervised Monocular Depth EstimationWenyao Zhang, Hongsi Liu, Bohan Li, Jiawei He 等ICCV 2025 · 被引用 2 次
- UniScene: Unified Occupancy-centric Driving Scene GenerationBohan Li, Jiazhe Guo, Hongsi Liu, Yingshuang Zou 等CVPR 2025
它引用的顶会 Paper19
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
- Hierarchical Neural Architecture Search for Deep Stereo MatchingXuelian Cheng, Yiran Zhong, Mehrtash Harandi, Yuchao Dai 等NeurIPS 2020 · 被引用 436 次
相关 Paper
- PatchScene: Patch-based Voxel Diffusion Model for Large-Scale Scene CompletionQingdong Xu, Jiajun Zhu, Shilin Zhu, Xinjing He 等CVPR 2026
- Cascaded Context Pyramid for Full-Resolution 3D Semantic Scene CompletionPingping Zhang, Wei Liu, Yinjie Lei, Huchuan Lu 等ICCV 2019 · 被引用 79 次
- Voxel Proposal Network via Multi-Frame Knowledge Distillation for Semantic Scene CompletionLubo Wang, Di Lin, Kairui Yang, Ruonan Liu 等NeurIPS 2024 · 被引用 14 次
- 3D Shape Generation and Completion through Point-Voxel DiffusionLinqi Zhou, Yilun Du, Jiajun WuICCV 2021 · 被引用 681 次
- HD²-SSC: High-Dimension High-Density Semantic Scene Completion for Autonomous DrivingZhiwen Yang, Yuxin PengAAAI 2026
