InstaGen: Enhancing Object Detection by Training on Synthetic Dataset
Chengjian Feng, Yujie Zhong, Zequn Jie, Weidi Xie, Lin Ma
2024年份
6顶会引用
摘要
https://fcjian.github.io/InstaGen Figure 1. (a) The synthetic images generated from Stable Diffusion and our proposed InstaGen, which can serve as a dataset synthesizer for sourcing photo-realistic images and instance bounding boxes at scale. (b) On open-vocabulary detection, training on synthetic images demonstrates significant improvement over CLIP-based methods on novel categories. (c) Training on the synthetic images generated from InstaGen also enhances the detection performance in close-set scenario, particularly in data-sparse circumstances.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- ODGEN: Domain-specific Object Detection Data Generation with Diffusion ModelsJingyuan Zhu, Shiyu Li, Yuxuan Liu, Jian Yuan 等NeurIPS 2024 · 被引用 32 次
- Sample-Efficient Multi-Round Generative Data Augmentation for Long-Tail Instance SegmentationByunghyun Kim, Minyoung Bae, Jae-Gil LeeNeurIPS 2025 · 被引用 3 次
- Gen4Track: A Tuning-free Data Augmentation Framework via Self-correcting Diffusion Model for Vision-Language TrackingJiawei Ge, Xinyu Zhang, Jiuxin Cao, Xuelin Zhu 等ACM MM 2025 · 被引用 3 次
- Advancing Visual Large Language Model for Multi-Granular Versatile PerceptionWentao Xiang, Haoxian Tan, Yujie Zhong, Cong Wei 等ICCV 2025 · 被引用 1 次
- On the Difficulty of Learning a Meta-network for Training Data SelectionZilin Du, Junqi Zhao, Albert Boyang LiICML 2026
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
相关 Paper
- Fake it Till You Make it: Learning Transferable Representations from Synthetic ImageNet ClonesMert Bülent Sariyildiz, Karteek Alahari, Diane Larlus, Yannis KalantidisCVPR 2023
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image GenerationCihang Peng, Qiming Hou, Zhong Ren, Kun ZhouICCV 2025
- FodFoM: Fake Outlier Data by Foundation Models Creates Stronger Visual Out-of-Distribution DetectorJiankang Chen, Ling Deng, Zhiyong Gan, Wei-Shi Zheng 等ACM MM 2024 · 被引用 3 次
- Generating Enhanced Negatives for Training Language-Based Object DetectorsShiyu Zhao, Long Zhao, Vijay Kumar B. G, Yumin Suh 等CVPR 2024 · 被引用 4 次
- ProxyDet: Synthesizing Proxy Novel Classes via Classwise Mixup for Open-Vocabulary Object DetectionJoonhyun Jeong, Geondo Park, Jayeon Yoo, Hyungsik Jung 等AAAI 2024 · 被引用 18 次
