Identity From Here, Pose From There: Self-Supervised Disentanglement and Generation of Objects Using Unlabeled Videos
Fanyi Xiao, Haotian Liu, Yong Jae Lee
Abstract
We propose a novel approach that disentangles the identity and pose of objects for image generation. Our model takes as input an ID image and a pose image, and generates an output image with the identity of the ID image and the pose of the pose image. Unlike most previous unsupervised work which rely on cyclic constraints, which can often be brittle, we instead propose to learn this in a self-supervised way. Specifically, we leverage unlabeled videos to automatically construct pseudo ground-truth targets to directly supervise our model. To enforce disentanglement, we propose a novel disentanglement loss, and to improve realism, we propose a pixel-verification loss in which the generated image's pixels must trace back to the ID input. We conduct extensive experiments on both synthetic and real images to demonstrate improved realism, diversity, and ID/pose disentanglement compared to existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eae947c8-8a08-4976-94a0-c76f5e2b7653Cited by top-tier papers3
- Neural Head Reenactment with Latent Pose DescriptorsEgor Burkov, Igor Pasechnik, Artur Grigorev, Victor S. LempitskyCVPR 2020
- MixNMatch: Multifactor Disentanglement and Encoding for Conditional Image GenerationYuheng Li, Krishna Kumar Singh, Utkarsh Ojha, Yong Jae LeeCVPR 2020
- Disentangled Image Generation Through Structured Noise InjectionYazeed Alharbi, Peter WonkaCVPR 2020
Related papers
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 37 citations
- Unsupervised Robust Disentangling of Latent Characteristics for Image SynthesisPatrick Esser, Johannes Haux, Björn OmmerICCV 2019 · 40 citations
- Self-Supervised 3D Human Pose Estimation via Part Guided Novel Image SynthesisJogendra Nath Kundu, Siddharth Seth, Varun Jampani, Mugalodi Rakesh et al.CVPR 2020
- MAPConNet: Self-supervised 3D Pose Transfer with Mesh and Point Contrastive LearningJiaze Sun, Zhixiang Chen, Tae-Kyun KimICCV 2023 · 2 citations
- Image Animation with Perturbed MasksYoav Shalev, Lior WolfCVPR 2022 · 6 citations
