Multimodal Structure-Consistent Image-to-Image Translation
Che-Tsung Lin, Yen-Yi Wu, Po-Hao Hsu, Shang-Hong Lai
Abstract
Unpaired image-to-image translation is proven quite effective in boosting a CNN-based object detector for a different domain by means of data augmentation that can well preserve the image-objects in the translated images. Recently, multimodal GAN (Generative Adversarial Network) models have been proposed and were expected to further boost the detector accuracy by generating a diverse collection of images in the target domain, given only a single/labelled image in the source domain. However, images generated by multimodal GANs would achieve even worse detection accuracy than the ones by a unimodal GAN with better object preservation. In this work, we introduce cycle-structure consistency for generating diverse and structure-preserved translated images across complex domains, such as between day and night, for object detector training. Qualitative results show that our model, Multimodal AugGAN, can generate diverse and realistic images for the target domain. For quantitative comparisons, we evaluate other competing methods and ours by using the generated images to train YOLO, Faster R-CNN and FCN models and prove that our model achieves significant improvement and outperforms other methods on the detection accuracies and the FCN scores. Also, we demonstrate that our model could provide more diverse object appearances in the target domain through comparison on the perceptual distance metric.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext abe3bfc5-76f8-4eb9-8586-d1b3c7b1c5c3Cited by top-tier papers4
- NightLab: A Dual-level Architecture with Hardness Detection for Segmentation at NightXueqing Deng, Peng Wang, Xiaochen Lian, Shawn D. NewsamCVPR 2022 · 51 citations
- Unsupervised Coherent Video Cartoonization with Perceptual Motion ConsistencyZhenhuan Liu, Liang Li, Huajie Jiang, Xin Jin et al.AAAI 2022 · 7 citations
- Dark Side Augmentation: Generating Diverse Night Examples for Metric LearningAlbert Mohwald, Tomás Jenícek, Ondrej ChumICCV 2023 · 7 citations
- CoMoGAN: Continuous Model-Guided Image-to-Image TranslationFabio Pizzati, Pietro Cerri, Raoul de CharetteCVPR 2021
Related papers
- GA-DAN: Geometry-Aware Domain Adaptation Network for Scene Text Detection and RecognitionFangneng Zhan, Chuhui Xue, Shijian LuICCV 2019 · 89 citations
- Benign Examples: Imperceptible Changes Can Enhance Image Translation PerformanceVignesh Srinivasan, Klaus-Robert Müller, Wojciech Samek, Shinichi NakajimaAAAI 2020 · 2 citations
- Learning to Transfer: Unsupervised Domain Translation via Meta-LearningJianxin Lin, Yijun Wang, Zhibo Chen, Tianyu HeAAAI 2020 · 10 citations
- On Translation and Reconstruction Guarantees of the Cycle-Consistent Generative Adversarial NetworksAnish Chakrabarty, Swagatam DasNeurIPS 2022 · 5 citations
- Breaking the Dilemma of Medical Image-to-image TranslationLingke Kong, Chenyu Lian, Detian Huang, Zhenjiang Li et al.NeurIPS 2021 · 234 citations
