Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion
Wentao Qu, Guofeng Mei, Jing Wang, Yujiao Wu, Xiaoshui Huang, Liang Xiao
Abstract
Denoising Diffusion Probabilistic Models (DDPMs) have shown success in robust 3D object detection tasks. Existing methods often rely on the score matching from 3D boxes or pre-trained diffusion priors. However, they typically require multi-step iterations in inference, which limits efficiency. To address this, we propose a Robust single-stage fully Sparse 3D object Detection Network with a Detachable Latent Framework (DLF) of DDPMs, named RSDNet. Specifically, RSDNet learns the denoising process in latent feature spaces through lightweight denoising networks like multi-level denoising autoencoders (DAEs). This enables RSDNet to effectively understand scene distributions under multi-level perturbations, achieving robust and reliable detection. Meanwhile, we reformulate the noising and denoising mechanisms of DDPMs, enabling DLF to construct multi-type and multi-level noise samples and targets, enhancing RSDNet robustness to multiple perturbations. Furthermore, a semantic-geometric conditional guidance is introduced to perceive the object boundaries and shapes, alleviating the center feature missing problem in sparse representations, enabling RSDNet to perform in a fully sparse detection pipeline. Moreover, the detachable denoising network design of DLF enables RSDNet to perform single-step detection in inference, further enhancing detection efficiency. Extensive experiments on public benchmarks show that RSDNet can outperform existing methods, achieving state-of-the-art detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3cdc62e5-df58-44d8-9815-d625bf086e86Cited by top-tier papers3
- A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene GenerationWentao Qu, Guofeng Mei, Yang Wu, Yongshun Gong et al.CVPR 2026 · 4 citations
- GEM: Generating LiDAR World Model via Deformable MambaYang Wu, Zhaojiang Liu, Qiang Meng, Youquan Liu et al.CVPR 2026 · 1 citation
- Bézier Degradation Modeling for LiDAR-based Human Motion CaptureXiaoqi An, Lin Zhao, Jun Li, Chen Gong et al.CVPR 2026
Builds on28
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- STD: Sparse-to-Dense 3D Object Detector for Point CloudZetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen et al.ICCV 2019 · 840 citations
Related papers
- An End-to-End Robust Point Cloud Semantic Segmentation Network with Single-Step Conditional Diffusion ModelsWentao Qu, Jing Wang, Yongshun Gong, Xiaoshui Huang et al.CVPR 2025
- Sparse2Dense: Learning to Densify 3D Features for 3D Object DetectionTianyu Wang, Xiaowei Hu, Zhengzhe Liu, Chi-Wing FuNeurIPS 2022 · 24 citations
- RSN: Range Sparse Net for Efficient, Accurate LiDAR 3D Object DetectionPei Sun, Weiyue Wang, Yuning Chai, Gamaleldin Elsayed et al.CVPR 2021
- SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single ImagesZixuan Huang, Mark Boss, Aaryaman Vasishta, James M. Rehg et al.CVPR 2025
- Promptable 3-D Object Localization with Latent Diffusion ModelsCheng-Yao Hong, Li-Heng Wang, Tyng-Luh LiuNeurIPS 2025
