DUNIT: Detection-Based Unsupervised Image-to-Image Translation
Deblina Bhattacharjee, Seungryong Kim, Guillaume Vizier, Mathieu Salzmann
Abstract
Image-to-image translation has made great strides in recent years, with current techniques being able to handle unpaired training images and to account for the multimodality of the translation problem. Despite this, most methods treat the image as a whole, which makes the results they produce for content-rich scenes less realistic. In this paper, we introduce a Detection-based Unsupervised Image-to-image Translation (DUNIT) approach that explicitly accounts for the object instances in the translation process. To this end, we extract separate representations for the global image and for the instances, which we then fuse into a common representation from which we generate the translated image. This allows us to preserve the detailed content of object instances, while still modeling the fact that we aim to produce an image of a single consistent scene. We introduce an instance consistency loss to maintain the coherence between the detections. Furthermore, by incorporating a detector into our architecture, we can still exploit object instances at test time. As evidenced by our experiments, this allows us to outperform the state-of-the-art unsupervised image-to-image translation methods. Furthermore, our approach can also be used as an unsupervised domain adaptation strategy for object detection, and it also achieves state-of-the-art performance on this task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a884aff5-b999-403a-b9a8-eaa6ee70fa23Cited by top-tier papers15
- Frequency Domain Image Translation: More Photo-realistic, Better Identity-preservingMu Cai, Hong Zhang, Huijuan Huang, Qichuan Geng et al.ICCV 2021 · 118 citations
- Dual Path Learning for Domain Adaptation of Semantic SegmentationYiting Cheng, Fangyun Wei, Jianmin Bao, Dong Chen et al.ICCV 2021 · 77 citations
- MuIT: An End-to-End Multitask Learning TransformerDeblina Bhattacharjee, Tong Zhang, Sabine Süsstrunk, Mathieu SalzmannCVPR 2022 · 64 citations
- InstaFormer: Instance-Aware Image-to-Image Translation with TransformerSoohyun Kim, Jongbeom Baek, Jihye Park, Gyeongnyeon Kim et al.CVPR 2022 · 53 citations
- Few shot font generation via transferring similarity guided global style and quantization local styleWei Pan, Anna Zhu, Xinyu Zhou, Brian Kenji Iwana et al.ICCV 2023 · 24 citations
Related papers
- Rethinking the Truly Unsupervised Image-to-Image TranslationKyungjune Baek, Yunjey Choi, Youngjung Uh, Jaejun Yoo et al.ICCV 2021 · 115 citations
- Self-Supervised Dense Consistency Regularization for Image-to-Image TranslationMinsu Ko, Eunju Cha, Sungjoo Suh, Huijin Lee et al.CVPR 2022 · 25 citations
- Cross-Granularity Learning for Multi-Domain Image-to-Image TranslationHuiyuan Fu, Ting Yu, Xin Wang, Huadong MaACM MM 2020 · 4 citations
- Memory-Guided Unsupervised Image-to-Image TranslationSomi Jeong, Youngjung Kim, Eungbean Lee, Kwanghoon SohnCVPR 2021
- Unsupervised Image-to-Image Translation with Generative PriorShuai Yang, Liming Jiang, Ziwei Liu, Chen Change LoyCVPR 2022 · 51 citations
