Unbiased Object Detection Beyond Frequency with Visually Prompted Image Synthesis
Xinhao Cai, Liulei Li, Gensheng Pei, Tao Chen, Jinshan Pan, Yazhou Yao, Wenguan Wang
Abstract
This paper presents a generation-based debiasing framework for object detection. Prior debiasing methods are often limited by the representation diversity of samples, while naive generative augmentation often preserves the biases it aims to solve. Moreover, our analysis reveals that simply generating more data for rare classes is suboptimal due to two core issues: i) instance frequency is an incomplete proxy for the true data needs of a model, and ii) current layout-to-image synthesis lacks the fidelity and control to generate high-quality, complex scenes. To overcome this, we introduce the representation score (RS) to diagnose representational gaps beyond mere frequency, guiding the creation of new, unbiased layouts. To ensure high-quality synthesis, we replace ambiguous text prompts with a precise visual blueprint and employ a generative alignment strategy, which fosters communication between the detector and generator. Our method significantly narrows the performance gap for underrepresented object groups, e.g., improving large/rare instances by 4.4/3.6 mAP over the baseline, and surpassing prior L2I synthesis models by 15.9 mAP for layout accuracy in generated images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e996f58-e6a0-4bc7-9f23-6fad6846617cCited by top-tier papers3
- PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part SegmentationJianjian Yin, Tao Chen, Yi Chen, Gensheng Pei et al.CVPR 2026 · 6 citations
- Iris: Bringing Real-World Priors into Diffusion Model for Monocular Depth EstimationXinhao Cai, Gensheng Pei, Zeren Sun, Yazhou Yao et al.CVPR 2026 · 2 citations
- Learning 3D Representations for Spatial Intelligence from Unposed Multi-View ImagesBo Zhou, Qiuxia Lai, Zeren Sun, Xiangbo Shu et al.CVPR 2026 · 1 citation
Builds on36
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
Related papers
- Reconciling Object-Level and Global-Level Objectives for Long-Tail DetectionShaoyu Zhang, Chen Chen, Silong PengICCV 2023 · 9 citations
- Envisioning Beyond the Few: Disentangled Semantics and Primitives for Few-Shot Atypical Layout-to-Image GenerationNan Bao, Yifan Zhao, Wenzhuang Wang, Jia LiICML 2026 · 1 citation
- Learning Latent Concepts for Detecting Out-of-Distribution ObjectsTing Peng, Junhao Dong, Yew-Soon OngCVPR 2026
- Dual Data Alignment Makes AI-Generated Image Detector Easier GeneralizableRuoxin Chen, Junwei Xi, Zhiyuan Yan, Ke-Yue Zhang et al.NeurIPS 2025 · 78 citations
- A Plug-and-Play Method for Rare Human-Object Interactions Detection by Bridging Domain GapLijun Zhang, Wei Suo, Peng Wang, Yanning ZhangACM MM 2024 · 4 citations
