Unsupervised Robust Disentangling of Latent Characteristics for Image Synthesis
Patrick Esser, Johannes Haux, Björn Ommer
Abstract
Deep generative models come with the promise to learn an explainable representation for visual objects that allows image sampling, synthesis, and selective modification. The main challenge is to learn to properly model the independent latent characteristics of an object, especially its appearance and pose. We present a novel approach that learns disentangled representations of these characteristics and explains them individually. Training requires only pairs of images depicting the same object appearance, but no pose annotations. We propose an additional classifier that estimates the minimal amount of regularization required to enforce disentanglement. Thus both representations together can completely explain an image while being independent of each other. Previous methods based on adversarial approaches fail to enforce this independence, while methods based on variational approaches lead to uninformative representations. In experiments on diverse object categories, the approach successfully recombines pose and appearance to reconstruct and retarget novel synthesized images. We achieve significant improvements over state-of-the-art methods which utilize the same level of supervision, and reach performances comparable to those of pose-supervised approaches. However, we can handle the vast body of articulated object classes for which no pose models/annotations are available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0645a43f-e305-49dc-8f2b-23d31b33d371Cited by top-tier papers13
- Swapping Autoencoder for Deep Image ManipulationTaesung Park, Jun-Yan Zhu, Oliver Wang, Jingwan Lu et al.NeurIPS 2020 · 376 citations
- Content and Style Disentanglement for Artistic Style TransferDmytro Kotovenko, Artsiom Sanakoyeu, Sabine Lang, Björn OmmerICCV 2019 · 187 citations
- Network-to-Network Translation with Conditional Invertible Neural NetworksRobin Rombach, Patrick Esser, Björn OmmerNeurIPS 2020 · 49 citations
- Style Equalization: Unsupervised Learning of Controllable Generative Sequence ModelsJen-Hao Rick Chang, Ashish Shrivastava, Hema Koppula, Xiaoshuai Zhang et al.ICML 2022 · 21 citations
- Disentanglement Analysis with Partial Information DecompositionSeiya Tokui, Issei SatoICLR 2022 · 16 citations
Builds on1
Related papers
- MUST-GAN: Multi-Level Statistics Transfer for Self-Driven Person Image GenerationTianxiang Ma, Bo Peng, Wei Wang, Jing DongCVPR 2021
- Identity From Here, Pose From There: Self-Supervised Disentanglement and Generation of Objects Using Unlabeled VideosFanyi Xiao, Haotian Liu, Yong Jae LeeICCV 2019 · 16 citations
- Scaling-up Disentanglement for Image TranslationAviv Gabbay, Yedid HoshenICCV 2021 · 22 citations
- GIRAFFE: Representing Scenes As Compositional Generative Neural Feature FieldsMichael Niemeyer, Andreas GeigerCVPR 2021
- Motion Representations for Articulated AnimationAliaksandr Siarohin, Oliver J. Woodford, Jian Ren, Menglei Chai et al.CVPR 2021
