Scaling-up Disentanglement for Image Translation
Aviv Gabbay, Yedid Hoshen
Abstract
Image translation methods typically aim to manipulate a set of labeled attributes (given as supervision at training time e.g. domain label) while leaving the unlabeled attributes intact. Current methods achieve either: (i) disentanglement, which exhibits low visual fidelity and can only be satisfied where the attributes are perfectly uncorrelated. (ii) visually-plausible translations, which are clearly not disentangled. In this work, we propose OverLORD, a single framework for disentangling labeled and unlabeled attributes as well as synthesizing high-fidelity images, which is composed of two stages; (i) Disentanglement: Learning disentangled representations with latent optimization. Differently from previous approaches, we do not rely on adversarial training or any architectural biases. (ii) Synthesis: Training feed-forward encoders for inferring the learned attributes and tuning the generator in an adversarial manner to increase the perceptual quality. When the labeled and unlabeled attributes are correlated, we model an additional representation that accounts for the correlated attributes and improves disentanglement. We highlight that our flexible framework covers multiple settings as disentangling labeled attributes, pose and appearance, localized concepts, and shape and texture. We present significantly better disentanglement with higher translation quality and greater output diversity than state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- StyleAlign: Analysis and Applications of Aligned StyleGAN ModelsZongze Wu, Yotam Nitzan, Eli Shechtman, Dani LischinskiICLR 2022 · 65 citations
- General Image-to-Image Translation with One-Shot Image GuidanceBin Cheng, Zuhao Liu, Yunbo Peng, Yue LinICCV 2023 · 60 citations
- InstaFormer: Instance-Aware Image-to-Image Translation with TransformerSoohyun Kim, Jongbeom Baek, Jihye Park, Gyeongnyeon Kim et al.CVPR 2022 · 53 citations
- An Image is Worth More Than a Thousand Words: Towards Disentanglement in The WildAviv Gabbay, Niv Cohen, Yedid HoshenNeurIPS 2021 · 43 citations
- Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language ModelZipeng Xu, Tianwei Lin, Hao Tang, Fu Li et al.CVPR 2022 · 38 citations
Builds on9
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 956 citations
- Few-Shot Unsupervised Image-to-Image TranslationMing-Yu Liu, Xun Huang, Arun Mallya, Tero Karras et al.ICCV 2019 · 668 citations
- Swapping Autoencoder for Deep Image ManipulationTaesung Park, Jun-Yan Zhu, Oliver Wang, Jingwan Lu et al.NeurIPS 2020 · 376 citations
- Unsupervised Semantic Segmentation by Contrasting Object Mask ProposalsWouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis, Luc Van GoolICCV 2021 · 285 citations
- Only a matter of style: age transformation using a style-based regression modelYuval Alaluf, Or Patashnik, Daniel Cohen-OrSIGGRAPH 2021 · 143 citations
Related papers
- Demystifying Inter-Class DisentanglementAviv Gabbay, Yedid HoshenICLR 2020 · 63 citations
- Unsupervised Robust Disentangling of Latent Characteristics for Image SynthesisPatrick Esser, Johannes Haux, Björn OmmerICCV 2019 · 40 citations
- Harnessing the Conditioning Sensorium for Improved Image TranslationCooper Nederhood, Nicholas I. Kolkin, Deqing Fu, Jason SalavonICCV 2021 · 6 citations
- Identity From Here, Pose From There: Self-Supervised Disentanglement and Generation of Objects Using Unlabeled VideosFanyi Xiao, Haotian Liu, Yong Jae LeeICCV 2019 · 16 citations
- Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept PersonalizationTsai-Shien Chen, Aliaksandr Siarohin, Gordon Guocheng Qian, Kuan-Chieh Jackson Wang et al.CVPR 2026 · 4 citations
