Learning Realistic Human Reposing using Cyclic Self-Supervision with 3D Shape, Pose, and Appearance Consistency
Soubhik Sanyal, Betty J. Mohler, Alex Vorobiov, Larry Davis, Timo Bolkart, Javier Romero, Matthew Loper, Michael J. Black
Abstract
Synthesizing images of a person in novel poses from a single image is a highly ambiguous task. Most existing approaches require paired training images; i.e. images of the same person with the same clothing in different poses. However, obtaining sufficiently large datasets with paired data is challenging and costly. Previous methods that forego paired supervision lack realism. We propose a self-supervised framework named SPICE (Self-supervised Person Image CrEation) that closes the image quality gap with supervised methods. The key insight enabling self-supervision is to exploit 3D information about the human body in several ways. First, the 3D body shape must remain unchanged when reposing. Second, representing body pose in 3D enables reasoning about self occlusions. Third, 3D body parts that are visible before and after reposing, should have similar appearance features. Once trained, SPICE takes an image of a person and generates a new image of that person in a new target pose. SPICE achieves state-of-the-art performance on the DeepFashion dataset, improving the FID score from 29.9 to 7.8 compared with previous unsupervised methods, and with performance similar to the state-of-the-art supervised method (6.4). SPICE also generates temporally coherent videos given an input image and a sequence of poses, despite being trained on static images only.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c817704-446d-4499-aaa8-59ccc0b9bdf8Cited by top-tier papers8
- HumanNeRF: Free-viewpoint Rendering of Moving People from Monocular VideoChung-Yi Weng, Brian Curless, Pratul P. Srinivasan, Jonathan T. Barron et al.CVPR 2022 · 411 citations
- InsetGAN for Full-Body Image GenerationAnna Frühstück, Krishna Kumar Singh, Eli Shechtman, Niloy J. Mitra et al.CVPR 2022 · 52 citations
- 3DHumanGAN: 3D-Aware Human Image Generation with 3D Pose MappingZhuoqian Yang, Shikai Li, Wayne Wu, Bo DaiICCV 2023 · 19 citations
- Collecting The Puzzle Pieces: Disentangled Self-Driven Human Pose Transfer by Permuting TexturesNannan Li, Kevin J. Shih, Bryan A. PlummerICCV 2023 · 9 citations
- BodyGAN: General-purpose Controllable Neural Human Body GenerationChaojie Yang, Hanhui Li, Shengjie Wu, Shengkai Zhang et al.CVPR 2022 · 8 citations
Builds on8
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View SynthesisWen Liu, Zhixin Piao, Jie Min, Wenhan Luo et al.ICCV 2019 · 285 citations
- Human Synthesis and Scene CompositingMihai Zanfir, Elisabeta Oneata, Alin-Ionut Popa, Andrei Zanfir et al.AAAI 2020 · 24 citations
- Structure-aware Person Image Generation with Pose Decomposition and Semantic CorrelationJilin Tang, Yi Yuan, Tianjia Shao, Yong Liu et al.AAAI 2021 · 22 citations
- Cross-Domain Correspondence Learning for Exemplar-Based Image TranslationPan Zhang, Bo Zhang, Dong Chen, Lu Yuan et al.CVPR 2020
Related papers
- MUST-GAN: Multi-Level Statistics Transfer for Self-Driven Person Image GenerationTianxiang Ma, Bo Peng, Wei Wang, Jing DongCVPR 2021
- Learning High Fidelity Depths of Dressed Humans by Watching Social Media Dance VideosYasamin Jafarian, Hyun Soo ParkCVPR 2021
- Self-supervised Correlation Mining Network for Person Image GenerationZijian Wang, Xingqun Qi, Kun Yuan, Muyi SunCVPR 2022 · 16 citations
- SelfPose3d: Self-Supervised Multi-Person Multi-View 3d Pose EstimationVinkle Srivastav, Keqi Chen, Nicolas PadoyCVPR 2024 · 17 citations
- Identity From Here, Pose From There: Self-Supervised Disentanglement and Generation of Objects Using Unlabeled VideosFanyi Xiao, Haotian Liu, Yong Jae LeeICCV 2019 · 16 citations
