Pixel Codec Avatars
Shugao Ma, Tomas Simon, Jason M. Saragih, Dawei Wang, Yuecheng Li, Fernando De la Torre, Yaser Sheikh
Abstract
Telecommunication with photorealistic avatars in virtual or augmented reality is a promising path for achieving authentic face-to-face communication in 3D over remote physical distances. In this work, we present the Pixel Codec Avatars (PiCA): a deep generative model of 3D human faces that achieves state of the art reconstruction performance while being computationally efficient and adaptive to the rendering conditions during execution. Our model combines two core ideas: (1) a fully convolutional architecture for decoding spatially varying features, and (2) a renderingadaptive per-pixel decoder. Both techniques are integrated via a dense surface representation that is learned in a weakly-supervised manner from low-topology mesh tracking over training images. We demonstrate that PiCA improves reconstruction over existing techniques across testing expressions and views on persons of different gender and skin tone. Importantly, we show that the PiCA model is much smaller than the state-of-art baseline model, and makes multi-person telecommunicaiton possible: on a single Oculus Quest 2 mobile VR headset, 5 avatars are rendered in realtime in the same scene.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3159c69a-9346-4f05-8dcd-836fab0c7b4bCited by top-tier papers61
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang et al.NeurIPS 2024 · 253 citations
- Neural Head Avatars from Monocular RGB VideosPhilip-William Grassal, Malte Prinzler, Titus Leistner, Carsten Rother et al.CVPR 2022 · 173 citations
- The Power of Points for Modeling Humans in ClothingQianli Ma, Jinlong Yang, Siyu Tang, Michael J. BlackICCV 2021 · 124 citations
- P2P: Tuning Pre-trained Image Models for Point Cloud Analysis with Point-to-Pixel PromptingZiyi Wang, Xumin Yu, Yongming Rao, Jie Zhou et al.NeurIPS 2022 · 121 citations
- SplattingAvatar: Realistic Real-Time Human Avatars With Mesh-Embedded Gaussian SplattingZhijing Shao, Zhaolong Wang, Zhuang Li, Duotun Wang et al.CVPR 2024 · 92 citations
Builds on5
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Multiview Neural Surface Reconstruction by Disentangling Geometry and AppearanceLior Yariv, Yoni Kasten, Dror Moran, Meirav Galun et al.NeurIPS 2020 · 1,010 citations
- The eyes have it: an integrated eye and face model for photorealistic facial animationGabriel Schwartz, Shih-En Wei, Te-Li Wang, Stephen Lombardi et al.SIGGRAPH 2020 · 54 citations
- A Decoupled 3D Facial Shape Model by Adversarial TrainingVictoria Fernández Abrevaya, Adnane Boukhayma, Stefanie Wuhrer, Edmond BoyerICCV 2019 · 36 citations
Related papers
- ESCA: Enabling Seamless Codec Avatar Execution through Algorithm and Hardware Co-Optimization for Virtual RealityMingzhi Zhu, Ding Shang, Sai Qian ZhangNeurIPS 2025
- Auto-CARD: Efficient and Robust Codec Avatar Driving for Real-time Mobile TelepresenceYonggan Fu, Yuecheng Li, Chenghui Li, Jason M. Saragih et al.CVPR 2023
- LUCAS: Layered Universal Codec AvatarsDi Liu, Teng Deng, Giljoo Nam, Yu Rong et al.CVPR 2025
- Pixel-Aligned Volumetric AvatarsAmit Raj, Michael Zollhöfer, Tomas Simon, Jason M. Saragih et al.CVPR 2021
- Universal Facial Encoding of Codec Avatars from VR HeadsetsShaojie Bai, Te-Li Wang, Chenghui Li, Akshay Venkatesh et al.SIGGRAPH 2024 · 4 citations
