Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining
Junxuan Li, Rawal Khirodkar, Egor Zakharov, Jihyun Lee, Zhaoen Su, Yuan Dong, Julieta Martinez, Kai Li, Qingyang Tan, Takaaki Shiratori, Matthew Hu, Peihong Guo
Abstract
High-quality 3D avatar modeling faces a critical tradeoff between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and the domain gap between the studio environment and the real world. On the other hand, recent large-scale avatar models trained on millions of in-the-wild samples show promise for generalization across a wide range of identities, yet the resulting avatars are often of low-quality due to inherent 3D ambiguities. To address this, we present Large-Scale Codec Avatars (LCA), a high-fidelity, full-body * Core contributors † Project lead 3D avatar model that generalizes to world-scale populations in a feedforward manner, enabling efficient inference. Inspired by the success of large language models and vision foundation models, we present, for the first time, a pre/posttraining paradigm for 3D avatar modeling at scale: we pretrain on 1M in-the-wild videos to learn broad priors over appearance and geometry, then post-train on high-quality curated data to enhance expressivity and fidelity. LCA generalizes across hair styles, clothing, and demographics while providing precise, fine-grained facial expressions and finger-level articulation control, with strong identity preservation. Notably, we observe emergent generalization to relightability and loose garment support to unconstrained inputs, and zero-shot robustness to stylized imagery, despite the absence of direct supervision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on32
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder et al.ICML 2024 · 513 citations
- Animatable Neural Radiance Fields for Modeling Dynamic Human BodiesSida Peng, Junting Dong, Qianqian Wang, Shangzhan Zhang et al.ICCV 2021 · 461 citations
Related papers
- LUCAS: Layered Universal Codec AvatarsDi Liu, Teng Deng, Giljoo Nam, Yu Rong et al.CVPR 2025
- Relightable Full-Body Gaussian Codec AvatarsShaofei Wang, Tomas Simon, Igor Santesteban, Timur M. Bagautdinov et al.SIGGRAPH 2025 · 6 citations
- GAIA: Zero-shot Talking Avatar GenerationTianyu He, Junliang Guo, Runyi Yu, Yuchi Wang et al.ICLR 2024 · 51 citations
- GAIA: Generative Animatable Interactive Avatars with Expression-conditioned GaussiansZhengming Yu, Tianye Li, Jingxiang Sun, Omer Shapira et al.SIGGRAPH 2025 · 2 citations
- Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal PriorChen Guo, Junxuan Li, Yash Kant, Yaser Sheikh et al.CVPR 2025
