DiffBEV: Conditional Diffusion Model for Bird's Eye View Perception
Jiayu Zou, Kun Tian, Zheng Zhu, Yun Ye, Xingang Wang
Abstract
BEV perception is of great importance in the field of autonomous driving, serving as the cornerstone of planning, controlling, and motion prediction. The quality of the BEV feature highly affects the performance of BEV perception. However, taking the noises in camera parameters and Li-DAR scans into consideration, we usually obtain BEV representation with harmful noises. Diffusion models naturally have the ability to denoise noisy samples to the ideal data, which motivates us to utilize the diffusion model to get a better BEV representation. In this work, we propose an endto-end framework, named DiffBEV, to exploit the potential of diffusion model to generate a more comprehensive BEV representation. To the best of our knowledge, we are the first to apply diffusion model to BEV perception. In practice, we design three types of conditions to guide the training of the diffusion model which denoises the coarse samples and refines the semantic feature in a progressive way. What's more, a cross-attention module is leveraged to fuse the context of BEV feature and the semantic content of conditional diffusion model. DiffBEV achieves a 25.9% mIoU on the nuScenes dataset, which is 6.2% higher than the bestperforming existing approach. Quantitative and qualitative results on multiple benchmarks demonstrate the effectiveness of DiffBEV in BEV semantic segmentation and 3D object detection tasks. The code ‡ will be available soon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5c98f01-1ce6-4e06-bae3-7c7a2b4ba34bCited by top-tier papers10
- MonoDiff: Monocular 3D Object Detection and Pose Estimation with Diffusion ModelsYasiru Ranasinghe, Deepti Hegde, Vishal M. PatelCVPR 2024 · 21 citations
- Improving Bird's Eye View Semantic Segmentation by Task DecompositionTianhao Zhao, Yongcan Chen, Yu Wu, Tianyang Liu et al.CVPR 2024 · 11 citations
- Pragmatic Heterogeneous Collaborative Perception via Generative Communication MechanismJunfei Zhou, Penglin Dai, Quanmin Wei, Bingyi Liu et al.NeurIPS 2025 · 11 citations
- VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector QuantizationYiwei Zhang, Jin Gao, Fudong Ge, Guan Luo et al.NeurIPS 2024 · 3 citations
- Intriguing Properties of Diffusion Models: An Empirical Study of the Natural Attack Capability in Text-to-Image Generative ModelsTakami Sato, Justin Yue, Nanze Chen, Ningfei Wang et al.CVPR 2024 · 2 citations
Builds on25
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
Related papers
- BEVDiffuser: Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth GuidanceXin Ye, Burhaneddin Yaman, Sheng Cheng, Feng Tao et al.CVPR 2025
- DDP: Diffusion Model for Dense Visual PredictionYuanfeng Ji, Zhe Chen, Enze Xie, Lanqing Hong et al.ICCV 2023 · 223 citations
- UniFusion: Unified Multi-view Fusion Transformer for Spatial-Temporal Representation in Bird's-Eye-ViewZequn Qin, Jingyu Chen, Chao Chen, Xiaozhi Chen et al.ICCV 2023 · 38 citations
- BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object DetectionGuowen Zhang, Chenhang He, Liyi Chen, Lei ZhangAAAI 2026 · 2 citations
- DriveEditor: A Unified 3D Information-Guided Framework for Controllable Object Editing in Driving ScenesYiyuan Liang, Zhiying Yan, Liqun Chen, Jiahuan Zhou et al.AAAI 2025 · 16 citations
