PIRenderer: Controllable Portrait Image Generation via Semantic Neural Rendering
Yurui Ren, Ge Li, Yuanqi Chen, Thomas H. Li, Shan Liu
Abstract
Generating portrait images by controlling the motions of existing faces is an important task of great consequence to social media industries. For easy use and intuitive control, semantically meaningful and fully disentangled parameters should be used as modifications. However, many existing techniques do not provide such fine-grained controls or use indirect editing methods i.e. mimic motions of other individuals. In this paper, a Portrait Image Neural Renderer (PIRenderer) is proposed to control the face motions with the parameters of three-dimensional morphable face models (3DMMs). The proposed model can generate photo-realistic portrait images with accurate movements according to intuitive modifications. Experiments on both direct and indirect editing tasks demonstrate the superiority of this model. Meanwhile, we further extend this model to tackle the audio-driven facial reenactment task by extracting sequential motions from audio inputs. We show that our model can generate coherent videos with convincing movements from only a single reference image and a driving audio stream. Our source code is available at https://github.com/RenYurui/PIRender .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 216e1d9d-1e2e-4c19-be59-4c0ae15f6c8cCited by top-tier papers81
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang et al.NeurIPS 2024 · 253 citations
- Generative Neural Articulated Radiance FieldsAlexander W. Bergman, Petr Kellnhofer, Wang Yifan, Eric R. Chan et al.NeurIPS 2022 · 144 citations
- StyleTalk: One-Shot Talking Head Generation with Controllable Speaking StylesYifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan et al.AAAI 2023 · 135 citations
- Generalizable and Animatable Gaussian Head AvatarXuangeng Chu, Tatsuya HaradaNeurIPS 2024 · 115 citations
- MagicAnimate: Temporally Consistent Human Image Animation using Diffusion ModelZhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan et al.CVPR 2024 · 106 citations
Builds on8
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 687 citations
- One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingTing-Chun Wang, Arun Mallya, Ming-Yu LiuCVPR 2021
- Neural Head Reenactment with Latent Pose DescriptorsEgor Burkov, Igor Pasechnik, Artur Grigorev, Victor S. LempitskyCVPR 2020
- StyleRig: Rigging StyleGAN for 3D Control Over Portrait ImagesAyush Tewari, Mohamed A. Elgharib, Gaurav Bharaj, Florian Bernard et al.CVPR 2020
Related papers
- PerformRecast: Expression and Head Pose Disentanglement for Portrait Video EditingJiadong Liang, Bojun Xiong, Jie Tian, Hua Li et al.CVPR 2026
- RigNeRF: Fully Controllable Neural 3D PortraitsShahRukh Athar, Zexiang Xu, Kalyan Sunkavalli, Eli Shechtman et al.CVPR 2022 · 117 citations
- DiffusionAvatars: Deferred Diffusion for High-fidelity 3D Head AvatarsTobias Kirschstein, Simon Giebenhain, Matthias NießnerCVPR 2024
- Single-Shot Implicit Morphable Faces with Consistent Texture ParameterizationConnor Z. Lin, Koki Nagano, Jan Kautz, Eric R. Chan et al.SIGGRAPH 2023 · 14 citations
- Neural Emotion Director: Speech-preserving semantic control of facial expressions in "in-the-wild" videosFoivos Paraperas Papantoniou, Panagiotis Paraskevas Filntisis, Petros Maragos, Anastasios RoussosCVPR 2022 · 32 citations
