T2Bs: Text-to-Character Blendshapes via Video Generation
Jiahao Luo, Chaoyang Wang, Michael Vasilkovsky, Vladislav Shakhrai, Di Liu, Peiye Zhuang, Sergey Tulyakov, Peter Wonka, Hsin-Ying Lee, James Davis, Jian Wang
Abstract
We present T2Bs, a framework for generating high-quality, animatable character head morphable models from text by combining static text-to-3D generation with video diffusion. Text-to-3D models produce detailed static geometry but lack motion synthesis, while video diffusion models generate motion with temporal and multi-view geometric inconsistencies. T2Bs bridges this gap by leveraging deformable 3D Gaussian splatting to align static 3D assets with video outputs. By constraining motion with static geometry and employing a view-dependent deformation MLP, T2Bs (i) outperforms existing 4D generation methods in accuracy and expressiveness while reducing video artifacts and view inconsistencies, and (ii) reconstructs smooth, coherent, fully registered 3D geometries designed to scale for building morphable models with diverse, realistic facial motions. This enables synthesizing expressive, animatable character heads that surpass current 4D generation techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on63
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
Related papers
- Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion ModelsHuan Ling, Seung Wook Kim, Antonio Torralba, Sanja Fidler et al.CVPR 2024
- Text-based Animatable 3D Avatars with Morphable Model AlignmentYiqian Wu, Malte Prinzler, Xiaogang Jin, Siyu TangSIGGRAPH 2025 · 1 citation
- Gaussian Variation Field Diffusion for High-Fidelity Video-to-4D SynthesisBowen Zhang, Sicheng Xu, Chuxin Wang, Jiaolong Yang et al.ICCV 2025 · 4 citations
- TeGA: Texture Space Gaussian Avatars for High-Resolution Dynamic Head ModelingGengyan Li, Paulo F. U. Gotardo, Timo Bolkart, Stephan J. Garbin et al.SIGGRAPH 2025 · 2 citations
- A Unified Approach for Text-and Image-Guided 4D Scene GenerationYufeng Zheng, Xueting Li, Koki Nagano, Sifei Liu et al.CVPR 2024 · 20 citations
