Artificial Dummies for Urban Dataset Augmentation
Antonín Vobecký, David Hurych, Michal Uricár, Patrick Pérez, Josef Sivic
Abstract
Existing datasets for training pedestrian detectors in images suffer from limited appearance and pose variation. The most challenging scenarios are rarely included because they are too difficult to capture due to safety reasons, or they are very unlikely to happen. The strict safety requirements in assisted and autonomous driving applications call for an extra high detection accuracy also in these rare situations. Having the ability to generate people images in arbitrary poses, with arbitrary appearances and embedded in different background scenes with varying illumination and weather conditions, is a crucial component for the development and testing of such applications. The contributions of this paper are three-fold. First, we describe an augmentation method for the controlled synthesis of urban scenes containing people, thus producing rare or never-seen situations. This is achieved with a data generator (called DummyNet) with disentangled control of the pose, the appearance, and the target background scene. Second, the proposed generator relies on novel network architecture and associated loss that takes into account the segmentation of the foreground person and its composition into the background scene. Finally, we demonstrate that the data generated by our DummyNet improve the performance of several existing person detectors across various datasets as well as in challenging situations, such as night-time conditions, where only a limited amount of training data is available. In the setup with only day-time data available, we improve the night-time detector by 17% log-average miss rate over the detector trained with the day-time data only.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- SinGAN: Learning a Generative Model From a Single Natural ImageTamar Rott Shaham, Tali Dekel, Tomer MichaeliICCV 2019 · 933 citations
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- Generative Modeling for Small-Data Object DetectionLanlan Liu, Michael Muelly, Jia Deng, Tomas Pfister et al.ICCV 2019 · 72 citations
- Human Synthesis and Scene CompositingMihai Zanfir, Elisabeta Oneata, Alin-Ionut Popa, Andrei Zanfir et al.AAAI 2020 · 24 citations
- Semi-Supervised Pedestrian Instance Synthesis and Detection With Mutual ReinforcementSi Wu, Sihao Lin, Wenhao Wu, Mohamed Azzam et al.ICCV 2019 · 8 citations
Related papers
- PedHunter: Occlusion Robust Pedestrian Detector in Crowded ScenesCheng Chi, Shifeng Zhang, Junliang Xing, Zhen Lei et al.AAAI 2020 · 118 citations
- PoseSyn: Synthesizing Diverse 3D Pose Data from In-the-Wild 2D DataChangHee Yang, Hyeonseop Song, Seokhun Choi, Seungwoo Lee et al.ICCV 2025 · 1 citation
- Generalizable Pedestrian Detection: The Elephant in the RoomIrtiza Hasan, Shengcai Liao, Jinpeng Li, Saad Ullah Akram et al.CVPR 2021
- Viperson: Flexibly Generating Virtual Identity for Person Re-IdentificationXiao-Wen Zhang, Delong Zhang, Yi-Xing Peng, Zhi Ouyang et al.ICCV 2025 · 2 citations
- DANNet: A One-Stage Domain Adaptation Network for Unsupervised Nighttime Semantic SegmentationXinyi Wu, Zhenyao Wu, Hao Guo, Lili Ju et al.CVPR 2021
