Harnessing the Conditioning Sensorium for Improved Image Translation
Cooper Nederhood, Nicholas I. Kolkin, Deqing Fu, Jason Salavon
摘要
Multi-modal domain translation typically refers to synthesizing a novel image that inherits certain localized attributes from a ‘content’ image (e.g. layout, semantics, or geometry), and inherits everything else (e.g. texture, lighting, sometimes even semantics) from a ‘style’ image. The dominant approach to this task is attempting to learn disentangled ‘content’ and ‘style’ representations from scratch. However, this is not only challenging, but ill-posed, as what users wish to preserve during translation varies depending on their goals. Motivated by this inherent ambiguity, we define ‘content’ based on conditioning information extracted by off-the-shelf pre-trained models. We then train our style extractor and image decoder with an easy to optimize set of reconstruction objectives. The wide variety of high-quality pre-trained models available and simple training procedure makes our approach straightforward to apply across numerous domains and definitions of ‘content’. Additionally it offers intuitive control over which aspects of ’content’ are preserved across domains. We evaluate our method on traditional, well-aligned, datasets such as CelebA-HQ, and propose two novel datasets for evaluation on more complex scenes: ClassicTV and FFHQ-Wild. Our approach, Sensorium, enables higher quality domain translation for more complex scenes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 被引用 840 次
- Guided Image-to-Image Translation With Bi-Directional Feature TransformationBadour Albahar, Jia-Bin HuangICCV 2019 · 被引用 102 次
- Make a Face: Towards Arbitrary High Fidelity Face ManipulationShengju Qian, Kwan-Yee Lin, Wayne Wu, Yangxiaokang Liu 等ICCV 2019 · 被引用 75 次
- Cross-Domain Correspondence Learning for Exemplar-Based Image TranslationPan Zhang, Bo Zhang, Dong Chen, Lu Yuan 等CVPR 2020
- Encoding in Style: A StyleGAN Encoder for Image-to-Image TranslationElad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan 等CVPR 2021
相关 Paper
- StEP: Style-Based Encoder Pre-Training for Multi-Modal Image SynthesisMoustafa Meshry, Yixuan Ren, Larry S. Davis, Abhinav ShrivastavaCVPR 2021
- Retrieval Guided Unsupervised Multi-domain Image to Image TranslationRaul Gomez, Yahui Liu, Marco De Nadai, Dimosthenis Karatzas 等ACM MM 2020 · 被引用 7 次
- Style-Guided and Disentangled Representation for Robust Image-to-Image TranslationJaewoong Choi, Dae Ha Kim, Byung Cheol SongAAAI 2022 · 被引用 9 次
- Scaling-up Disentanglement for Image TranslationAviv Gabbay, Yedid HoshenICCV 2021 · 被引用 22 次
- DINO: A Conditional Energy-Based GAN for Domain TranslationKonstantinos Vougioukas, Stavros Petridis, Maja PanticICLR 2021 · 被引用 8 次
