MMHead: Towards Fine-grained Multi-modal 3D Facial Animation
Sijing Wu, Yunhao Li, Yichao Yan, Huiyu Duan, Ziwei Liu, Guangtao Zhai
Abstract
The portrait talks about something. <emotion> The portrait looks surprised throughout the video clip.
The portrait in the video clip starts to frown, then sneers, wags its head, and finally starts talking. <emotion> The portrait's expression remains contemptuous throughout the video clip.
<expression> The subject begins with a strong inner brow raise and outer brow raise, with the lips parting widely and the jaw dropping slightly. As the sequence progresses, the inner brow raise remains consistent, . . . Towards the end of the sequence, the subject displays a mixture of facial movements, including a mouth stretch and tongue show. <head pose> The portrait maintains a steady head pose throughout the video, with no noticeable movements or rotations detected. <scenarios> 1. Finding out he won the lottery; 2. Witnessing a loved one's unexpected return; 3. ... <expression> The lips initially part with a high intensity, followed by the nostrils compressing slightly. ... As time progresses, the lid tightens, the lips continue to part, and the inner brows maintain a high level of intensity ...The nostrils compress periodically, and the cheek on the left side raises slightly. <head pose> The portrait turns its head to the right ... The head then switches direction, turning left gradually while also tilting slightly to the right. ... The head keeps moving upwards, turning left consistently at an increasing angle. Finally, the head settles in an upturned position with a slight leftward turn remaining. <scenarios> 1. The portrait is scrolling through social media and seeing posts from people he Dislike; 2. The portrait is watching a political debate where their least favorite candidate is speaking; 3. ... Figure 1: We present MMHead, the first multi-modal 3D facial animation dataset with hierarchical text annotations including abstract descriptions for overall actions and emotions, and fine-grained descriptions for expressions, head poses, as well as possible scenarios that may cause such emotions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46f0a593-5930-4a8b-97f5-89b955424562Cited by top-tier papers3
- UniLS: End-to-End Audio-Driven Avatars for Unified Listening and SpeakingXuangeng Chu, Ruicong Liu, Yifei Huang, Yun Liu et al.CVPR 2026 · 12 citations
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality AlignmentBin Wang, Yang Xu, Huan Zhao, Hao Zhang et al.ACM MM 2025
Builds on39
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot et al.ICML 2023 · 751 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
Related papers
- Flow-Guided One-Shot Talking Face Generation With a High-Resolution Audio-Visual DatasetZhimeng Zhang, Lincheng Li, Yu Ding, Changjie FanCVPR 2021
- ECAvatar: 3D Avatar Facial Animation with Controllable Identity and EmotionMinjing Yu, Delong Pang, Ziwen Kang, Zhiyao Sun et al.ACM MM 2024 · 4 citations
- VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive InteractionShiying Li, Xingqun Qi, Bingkun Yang, Weile Chen et al.AAAI 2026 · 2 citations
- MVPortrait: Text-Guided Motion and Emotion Control for Multi-view Vivid Portrait AnimationYukang Lin, Hokit Fung, Jianjin Xu, Zeping Ren et al.CVPR 2025
- Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained ControlHejia Chen, Haoxian Zhang, Shoulong Zhang, Xiaoqiang Liu et al.ICLR 2025
