LANIT: Language-Driven Image-to-Image Translation for Unlabeled Data
Jihye Park, Sunwoo Kim, Soohyun Kim, Seokju Cho, Jaejun Yoo, Youngjung Uh, Seungryong Kim
摘要
Existing techniques for image-to-image translation commonly have suffered from two critical problems: heavy reliance on per-sample domain annotation and/or inability to handle multiple attributes per image. Recent trulyunsupervised methods adopt clustering approaches to easily provide per-sample one-hot domain labels. However, they cannot account for the real-world setting: one sample may have multiple attributes. In addition, the semantics of the clusters are not easily coupled to human understanding. To overcome these, we present LANguage-driven Image-toimage Translation model, dubbed LANIT. We leverage easyto-obtain candidate attributes given in texts for a dataset: the similarity between images and attributes indicates persample domain labels. This formulation naturally enables multi-hot labels so that users can specify the target domain with a set of attributes in language. To account for the case that the initial prompts are inaccurate, we also present prompt learning. We further present domain regularization loss that enforces translated images to be mapped to the corresponding domain. Experiments on several standard benchmarks demonstrate that LANIT achieves comparable or superior performance to existing models. The code is available at github.com/KU-CVLAB/LANIT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
相关 Paper
- Rethinking the Truly Unsupervised Image-to-Image TranslationKyungjune Baek, Yunjey Choi, Youngjung Uh, Jaejun Yoo 等ICCV 2021 · 被引用 115 次
- Prompt Refinement with Image Pivot for Text-to-Image GenerationJingtao Zhan, Qingyao Ai, Yiqun Liu, Yingwei Pan 等ACL 2024
- Retrieval Guided Unsupervised Multi-domain Image to Image TranslationRaul Gomez, Yahui Liu, Marco De Nadai, Dimosthenis Karatzas 等ACM MM 2020 · 被引用 7 次
- Self-Supervised Dense Consistency Regularization for Image-to-Image TranslationMinsu Ko, Eunju Cha, Sungjoo Suh, Huijin Lee 等CVPR 2022 · 被引用 25 次
- Unsupervised Image-to-Image Translation with Generative PriorShuai Yang, Liming Jiang, Ziwei Liu, Chen Change LoyCVPR 2022 · 被引用 51 次
