PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment
Chaonan Ji, Jinwei Qi, Sheng Xu, Peng Zhang, Bang Zhang
摘要
Existing facial reenactment methods struggle with a trade-off between expressiveness and fine-grained controllability. Holistic facial reenactment models often sacrifice granular control for expressiveness, while methods designed for control may struggle with fidelity and robust disentanglement. Instead of treating facial motion as a monolithic signal, we explore an alternative compositional perspective. In this paper, we introduce PortraitDirector, a novel framework that formulates face reenactment as a hierarchical composition task, achieving high-fidelity and controllable results. We employ a Hierarchical Motion Disentanglement and Composition strategy, deconstructing facial motion into a Spatial Layer for physical movements and a Semantic Layer for emotional content. The Spatial Layer comprises: (i) global head pose, managed via a dedicated representation and injection pathway; (ii) spatially separated local facial expressions, distilled from cropped facial regions and purged of emotional cues via Emotion-Filtering Module leveraging an information bottleneck. The Semantic Layer contains a derived global emotion. The disentangled components are then recomposed into an expressive motion latent. Furthermore, we engineer the framework for real-time performance through a suite of optimizations, including diffusion distillation, causal attention and VAE acceleration. PortraitDirector achieves streaming, high-fidelity, controllable 512 x 512 face reenactment at 20 FPS with a end-to-end 800 ms latency on a single 5090 GPU.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper38
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou 等NeurIPS 2025 · 被引用 628 次
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu 等ICCV 2021 · 被引用 510 次
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang 等NeurIPS 2024 · 被引用 253 次
- EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion ModelXinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu 等SIGGRAPH 2022 · 被引用 150 次
相关 Paper
- DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion RepresentationsYuxiang Shi, Zhe Li, Yanwen Wang, Hao Zhu 等CVPR 2026 · 被引用 3 次
- Emotion-Conditioned Motion Sub-spaces with Flow Matching for Real-Time Audio-Driven Talking HeadsHaoyu Wang, Xiaozhe Xin, Xiaoyu Qin, Meiguang Jin 等AAAI 2026 · 被引用 1 次
- PersonaLive! Expressive Portrait Image Animation for Live StreamingZhiyuan Li, Chi-Man Pun, Chen Fang, Jue Wang 等CVPR 2026 · 被引用 6 次
- MVPortrait: Text-Guided Motion and Emotion Control for Multi-view Vivid Portrait AnimationYukang Lin, Hokit Fung, Jianjin Xu, Zeping Ren 等CVPR 2025
- Neural Emotion Director: Speech-preserving semantic control of facial expressions in "in-the-wild" videosFoivos Paraperas Papantoniou, Panagiotis Paraskevas Filntisis, Petros Maragos, Anastasios RoussosCVPR 2022 · 被引用 32 次
