GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
Kai Chen, Enze Xie, Zhe Chen, Yibo Wang, Lanqing Hong, Zhenguo Li, Dit-Yan Yeung
摘要
Diffusion models have attracted significant attention due to the remarkable ability to create content and generate data for tasks like image classification. However, the usage of diffusion models to generate the high-quality object detection data remains an underexplored area, where not only image-level perceptual quality but also geometric conditions such as bounding boxes and camera views are essential. Previous studies have utilized either copy-paste synthesis or layout-to-image (L2I) generation with specifically designed modules to encode the semantic layouts. In this paper, we propose the GeoDiffusion, a simple framework that can flexibly translate various geometric conditions into text prompts and empower pre-trained text-to-image (T2I) diffusion models for high-quality detection data generation. Unlike previous L2I methods, our GeoDiffusion is able to encode not only the bounding boxes but also extra geometric conditions such as camera views in self-driving scenes. Extensive experiments demonstrate GeoDiffusion outperforms previous L2I methods while maintaining 4x training time faster. To the best of our knowledge, this is the first work to adopt diffusion models for layout-to-image generation with geometric conditions and demonstrate that L2I-generated images can be beneficial for improving the performance of object detectors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityShenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta 等NeurIPS 2024 · 被引用 403 次
- Token Merging for Training-Free Semantic Binding in Text-to-Image SynthesisTaihang Hu, Linxuan Li, Joost van de Weijer, Hongcheng Gao 等NeurIPS 2024 · 被引用 45 次
- ODGEN: Domain-specific Object Detection Data Generation with Diffusion ModelsJingyuan Zhu, Shiyu Li, Yuxuan Liu, Jian Yuan 等NeurIPS 2024 · 被引用 32 次
- DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and PerceptionYibo Wang, Ruiyuan Gao, Kai Chen, Kaiqiang Zhou 等CVPR 2024 · 被引用 14 次
- Unbiased Object Detection Beyond Frequency with Visually Prompted Image SynthesisXinhao Cai, Liulei Li, Gensheng Pei, Tao Chen 等ICLR 2026 · 被引用 7 次
它引用的顶会 Paper26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
相关 Paper
- DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion GenerationHongbin Lin, Zilu Guo, Yifan Zhang, Shuaicheng Niu 等CVPR 2025
- ChangeDiff: A Multi-Temporal Change Detection Data Generator with Flexible Text Prompts via Diffusion ModelQi Zang, Jiayi Yang, Shuang Wang, Dong Zhao 等AAAI 2025 · 被引用 2 次
- Cycle-Consistent Learning for Joint Layout-to-Image Generation and Object DetectionXinhao Cai, Qiuxia Lai, Gensheng Pei, Xiangbo Shu 等ICCV 2025 · 被引用 1 次
- Visual Prototype Conditioned Focal Region Generation for UAV-Based Object DetectionWenhao Li, Zimeng Wu, Yu Wu, Zehua Fu 等CVPR 2026
- LAW-Diffusion: Complex Scene Generation by Diffusion with LayoutsBinbin Yang, Yi Luo, Ziliang Chen, Guangrun Wang 等ICCV 2023 · 被引用 21 次
