3DiffTection: 3D Object Detection with Geometry-Aware Diffusion Features
Chenfeng Xu, Huan Ling, Sanja Fidler, Or Litany
Abstract
3DiffTection introduces a novel method for 3D object detection from single images, utilizing a 3D-aware diffusion model for feature extraction. Addressing the resource-intensive nature of annotating large-scale 3D image data, our approach leverages pretrained diffusion models, traditionally used for 2D tasks, and adapts them for 3D detection through geometric and semantic tuning. Geometrically, we enhance the model to perform view synthesis from single images, incorporating an epipolar warp operator. This process utilizes easily accessible posed image data, eliminating the need for manual annotation. Semantically, the model is further refined on target detection data. Both stages utilize ControlNet, ensuring the preservation of original feature capabilities. Through our methodology, we obtain 3D-aware features that excel in identifying cross-view point correspondences. In 3D detection, 3DiffTection substantially surpasses previous benchmarks, e.g., Cube-RCNN, by 9.43% in AP3D on the Omni3D-ARkitscene dataset. Furthermore, 3DiffTection demonstrates robust label efficiency and generalizes well to cross-domain data, nearly matching fully-supervised models in zero-shot scenarios. Project page: https://research.nvidia.com/labs/toronto-ai/3difftection/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d947262b-fe92-4b0f-ac82-1795d4da1db4Cited by top-tier papers12
- Immiscible Diffusion: Accelerating Diffusion Training with Noise AssignmentYiheng Li, Heyang Jiang, Akio Kodaira, Masayoshi Tomizuka et al.NeurIPS 2024 · 19 citations
- Zero-to-Hero: Enhancing Zero-Shot Novel View Synthesis via Attention Map FilteringIdo Sobol, Chenfeng Xu, Or LitanyNeurIPS 2024 · 10 citations
- Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent DiffusionWentao Qu, Guofeng Mei, Jing Wang, Yujiao Wu et al.AAAI 2026 · 6 citations
- CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene GenerationLi Liang, Bo Miao, Xinyu Wang, Naveed Akhtar et al.NeurIPS 2025 · 4 citations
- CHARM3R: Towards Unseen Camera Height Robust Monocular 3D DetectorAbhinav Kumar, Yuliang Guo, Zhihao Zhang, Xinyu Huang et al.ICCV 2025 · 1 citation
Builds on33
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion PriorsSubin Jeon, In Cho, Junyoung Hong, Woong Oh Cho et al.CVPR 2026
- Large-Vocabulary 3D Diffusion Model with TransformerZiang Cao, Fangzhou Hong, Tong Wu, Liang Pan et al.ICLR 2024 · 54 citations
- MonoDiff: Monocular 3D Object Detection and Pose Estimation with Diffusion ModelsYasiru Ranasinghe, Deepti Hegde, Vishal M. PatelCVPR 2024 · 21 citations
- Generative Novel View Synthesis with 3D-Aware Diffusion ModelsEric R. Chan, Koki Nagano, Matthew A. Chan, Alexander W. Bergman et al.ICCV 2023 · 314 citations
- UniDet3D: Multi-dataset Indoor 3D Object DetectionMaksim Kolodiazhnyi, Anna Vorontsova, Matvey Skripkin, Danila Rukhovich et al.AAAI 2025 · 7 citations
