MMHead: Towards Fine-grained Multi-modal 3D Facial Animation
Sijing Wu, Yunhao Li, Yichao Yan, Huiyu Duan, Ziwei Liu, Guangtao Zhai
摘要
The portrait talks about something. <emotion> The portrait looks surprised throughout the video clip.
The portrait in the video clip starts to frown, then sneers, wags its head, and finally starts talking. <emotion> The portrait's expression remains contemptuous throughout the video clip.
<expression> The subject begins with a strong inner brow raise and outer brow raise, with the lips parting widely and the jaw dropping slightly. As the sequence progresses, the inner brow raise remains consistent, . . . Towards the end of the sequence, the subject displays a mixture of facial movements, including a mouth stretch and tongue show. <head pose> The portrait maintains a steady head pose throughout the video, with no noticeable movements or rotations detected. <scenarios> 1. Finding out he won the lottery; 2. Witnessing a loved one's unexpected return; 3. ... <expression> The lips initially part with a high intensity, followed by the nostrils compressing slightly. ... As time progresses, the lid tightens, the lips continue to part, and the inner brows maintain a high level of intensity ...The nostrils compress periodically, and the cheek on the left side raises slightly. <head pose> The portrait turns its head to the right ... The head then switches direction, turning left gradually while also tilting slightly to the right. ... The head keeps moving upwards, turning left consistently at an increasing angle. Finally, the head settles in an upturned position with a slight leftward turn remaining. <scenarios> 1. The portrait is scrolling through social media and seeing posts from people he Dislike; 2. The portrait is watching a political debate where their least favorite candidate is speaking; 3. ... Figure 1: We present MMHead, the first multi-modal 3D facial animation dataset with hierarchical text annotations including abstract descriptions for overall actions and emotions, and fine-grained descriptions for expressions, head poses, as well as possible scenarios that may cause such emotions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- UniLS: End-to-End Audio-Driven Avatars for Unified Listening and SpeakingXuangeng Chu, Ruicong Liu, Yifei Huang, Yun Liu 等CVPR 2026 · 被引用 12 次
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja 等ACM MM 2025 · 被引用 3 次
- PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality AlignmentBin Wang, Yang Xu, Huan Zhao, Hao Zhang 等ACM MM 2025
它引用的顶会 Paper39
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot 等ICML 2023 · 被引用 751 次
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu 等NeurIPS 2023 · 被引用 698 次
相关 Paper
- Flow-Guided One-Shot Talking Face Generation With a High-Resolution Audio-Visual DatasetZhimeng Zhang, Lincheng Li, Yu Ding, Changjie FanCVPR 2021
- ECAvatar: 3D Avatar Facial Animation with Controllable Identity and EmotionMinjing Yu, Delong Pang, Ziwen Kang, Zhiyao Sun 等ACM MM 2024 · 被引用 4 次
- VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive InteractionShiying Li, Xingqun Qi, Bingkun Yang, Weile Chen 等AAAI 2026 · 被引用 2 次
- MVPortrait: Text-Guided Motion and Emotion Control for Multi-view Vivid Portrait AnimationYukang Lin, Hokit Fung, Jianjin Xu, Zeping Ren 等CVPR 2025
- Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained ControlHejia Chen, Haoxian Zhang, Shoulong Zhang, Xiaoqiang Liu 等ICLR 2025
