CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models
Felix Taubner, Ruihang Zhang, Mathieu Tuli, David B. Lindell
摘要
Reconstructing photorealistic and dynamic portrait avatars from images is essential to many applications, including advertising, visual effects, and virtual reality. Depending on the application, avatar reconstruction involves different capture setups and constraints-for example, visual effects studios use camera arrays to capture hundreds of reference images, while content creators may seek to animate a single portrait image downloaded from the internet. As such, there is a large and heterogeneous ecosystem of methods for avatar reconstruction. Techniques based on multi-view stereo or neural rendering achieve the highest quality results, but require hundreds of reference images. Recent generative models produce convincing avatars from a single reference image, but visual fidelity lags behind multi-view techniques. Here, we present CAP4D: an approach that uses a morphable multi-view diffusion model to reconstruct photoreal 4D (dynamic 3D) portrait avatars from any number of reference images (i.e., one to 100) and animate and render them in real time. Our approach demonstrates stateof-the-art performance for single-, few-, and multi-image 4D portrait avatar reconstruction, and takes steps to bridge the gap in visual fidelity between single-image and multiview reconstruction techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face ReconstructionSimon Giebenhain, Tobias Kirschstein, Martin Rünz, Lourdes Agapito 等ICLR 2026 · 被引用 24 次
- FlexAvatar: Learning Complete 3D Head Avatars with Partial SupervisionTobias Kirschstein, Simon Giebenhain, Matthias NießnerCVPR 2026 · 被引用 10 次
- Avat3r: Large Animatable Gaussian Reconstruction Model for High-Fidelity 3D Head AvatarsTobias Kirschstein, Javier Romero, Artem Sevastopolsky, Matthias Nießner 等ICCV 2025 · 被引用 10 次
- UIKA: Fast Universal Head Avatar from Pose-Free ImagesZijian Wu, Boyao Zhou, Liangxiao Hu, Hongyu Liu 等CVPR 2026 · 被引用 6 次
- FastGHA: Generalized Few-Shot 3D Gaussian Head Avatars with Real-Time AnimationXinya Ji, Sebastian Weiss, Manuel Kansy, Jacek Naruniec 等ICLR 2026 · 被引用 6 次
它引用的顶会 Paper66
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- GeoDiff4D: Geometry-Aware Diffusion for 4D Head Avatar ReconstructionChao Xu, Xiaochen Zhao, Xiang Deng, Jingxiang Sun 等CVPR 2026
- Generalizable and Animatable Gaussian Head AvatarXuangeng Chu, Tatsuya HaradaNeurIPS 2024 · 被引用 115 次
- FaceCraft4D: Animated 3D Facial Avatar Generation from a Single ImageFei Yin, Mallikarjun B. R., Chun-Han Yao, Rafal K. Mantiuk 等ICCV 2025 · 被引用 2 次
- AniGS: Animatable Gaussian Avatar from a Single Image with Inconsistent Gaussian ReconstructionLingteng Qiu, Shenhao Zhu, Qi Zuo, Xiaodong Gu 等CVPR 2025
- Morphable Diffusion: 3D-Consistent Diffusion for Single-image Avatar CreationXiyi Chen, Marko Mihajlovic, Shaofei Wang, Sergey Prokudin 等CVPR 2024
