Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view Generation
Aviral Chharia, Fernando De la Torre
Abstract
High-fidelity 3D Gaussian head avatar generation is critical for applications such as AR/VR, telepresence, and digital humans. Existing methods depend on multi-view datasets, 3D captures, or intermediate 2D view synthesis. In contrast, we learn both conditional and unconditional 3D head models from randomly sampled 2D images alone, without using multi-view data, 3D supervision, or intermediate view generation. We introduce MVCHead, a single-shot state space model that enforces multi-view consistency (MVC) directly in the 3D representation while regressing 3D Gaussians under these constraints. At its core, we propose a Hierarchical State Space (HiSS) block that progressively refines Gaussians from coarse to fine, while capturing long-range dependencies. Within each HiSS block, we modify Mamba's standard unidirectional scan with the proposed Hierarchical Bi-directional State Scan (HiBiSS) that aligns recurrence with the axes along which multi-view inconsistencies are strongest. Finally, we design an SE(3) Multi-view Critic that judges whether a set of self-renders arises from a single underlying 3D configuration, rewarding cross-view pixel alignment without observing real multi-view pairs. MVCHead achieves state-of-the-art perceptual quality, surpasses prior methods in both texture and geometric consistency, and maintains comparable shape consistency. To demonstrate scalability, we release FaceGS-10K, the first large-scale dataset of ready-to-use 3D Gaussian head assets for training and evaluation of 3D head models. Project Page and code: https://humansensinglab.github.io/MVCHead/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64d4c323-c4e8-4714-b705-20bcc43bfe88Builds on53
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
Related papers
- High-Quality Full-Head 3D Avatar Generation from Any Single Portrait ImageYujie Gao, Chencheng Wang, Xianbing Sun, Jiahui Zhan et al.AAAI 2026
- MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar ReconstructionZijian Dong, Longteng Duan, Jie Song, Michael J. Black et al.ICCV 2025 · 4 citations
- MVGamba: Unify 3D Content Generation as State Space Sequence ModelingXuanyu Yi, Zike Wu, Qiuhong Shen, Qingshan Xu et al.NeurIPS 2024 · 27 citations
- SIGMAN: Scaling 3D Human Gaussian Generation with Millions of AssetsYuhang Yang, Fengqi Liu, Yixing Lu, Qin Zhao et al.ICCV 2025 · 5 citations
- FastGHA: Generalized Few-Shot 3D Gaussian Head Avatars with Real-Time AnimationXinya Ji, Sebastian Weiss, Manuel Kansy, Jacek Naruniec et al.ICLR 2026 · 6 citations
