Retrieval Guided Unsupervised Multi-domain Image to Image Translation
Raul Gomez, Yahui Liu, Marco De Nadai, Dimosthenis Karatzas, Bruno Lepri, Nicu Sebe
Abstract
Image to image translation aims to learn a mapping that transforms an image from one visual domain to another. Recent works assume that images descriptors can be disentangled into a domain-invariant content representation and a domain-specific style representation. Thus, translation models seek to preserve the content of source images while changing the style to a target visual domain. However, synthesizing new images is extremely challenging especially in multi-domain translations, as the network has to compose content and style to generate reliable and diverse images in multiple domains. In this paper we propose the use of an image retrieval system to assist the image-to-image translation task. First, we train an image-to-image translation model to map images to multiple domains. Then, we train an image retrieval model using real and generated images to find images similar to a query one in content but in a different domain. Finally, we exploit the image retrieval system to fine-tune the image-to-image translation model and generate higher quality images. Our experiments show the effectiveness of the proposed solution and highlight the contribution of the retrieval network, which can benefit from additional unlabeled data and help image-to-image translation models in the presence of scarce data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on2
Related papers
- Style-Guided and Disentangled Representation for Robust Image-to-Image TranslationJaewoong Choi, Dae Ha Kim, Byung Cheol SongAAAI 2022 · 9 citations
- Memory-Guided Unsupervised Image-to-Image TranslationSomi Jeong, Youngjung Kim, Eungbean Lee, Kwanghoon SohnCVPR 2021
- Harnessing the Conditioning Sensorium for Improved Image TranslationCooper Nederhood, Nicholas I. Kolkin, Deqing Fu, Jason SalavonICCV 2021 · 6 citations
- StEP: Style-Based Encoder Pre-Training for Multi-Modal Image SynthesisMoustafa Meshry, Yixuan Ren, Larry S. Davis, Abhinav ShrivastavaCVPR 2021
- Smoothing the Disentangled Latent Style Space for Unsupervised Image-to-Image TranslationYahui Liu, Enver Sangineto, Yajing Chen, Linchao Bao et al.CVPR 2021
