Smooth Online Multiple Appropriate Facial Reaction Generation
Weicheng Xie, Chunlin Yan, Siyang Song, Zitong Yu, Linlin Shen, Laizhong Cui
Abstract
In dyadic interactions, facial reactions are crucial for conveying an individuals' responses to their conversational partners. Individuals may exhibit varied but appropriate facial reactions (AFRs) when perceiving the same behavioral expression. Although some recent methods can already respond multiple appropriate facial reactions to the given human speaker behaviors, the AFRs generated by these methods often fail to adequately preserve crucial head motions, leading to visual jitter and unnatural transitions between generated AFR segments. In this paper, we propose a novel and generic PFLPosNet framework which addresses the aforementioned problems at both pre-processing and post-processing stages, where a new pose-aware face behavior localization method PFL is introduced to retain the head pose displacement information from the source data. In addition, the framework proposes a real-time head pose adjustment method, PosNet, to ensure continuity and smoothness in the visual output of the model when using data with correct head pose displacement. Experimental results demonstrate that our approach not only generates more coherent and natural facial reaction sequences but also significantly outperforms existing online MAFRG methods in terms of continuity and smoothness. Our code is made available at https://github.com/rainforcetime/PFLPosNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu et al.ICCV 2021 · 510 citations
- PIRenderer: Controllable Portrait Image Generation via Semantic Neural RenderingYurui Ren, Ge Li, Yuanqi Chen, Thomas H. Li et al.ICCV 2021 · 284 citations
- One-Shot Talking Face Generation from Single-Speaker Audio-Visual Correlation LearningSuzhen Wang, Lincheng Li, Yu Ding, Xin YuAAAI 2022 · 142 citations
- BeLFusion: Latent Diffusion for Behavior-Driven Human Motion PredictionGermán Barquero, Sergio Escalera, Cristina PalmeroICCV 2023 · 107 citations
- FaceVerse: a Fine-grained and Detail-controllable 3D Face Morphable Model from a Hybrid DatasetLizhen Wang, Zhiyuan Chen, Tao Yu, Chenguang Ma et al.CVPR 2022 · 87 citations
Related papers
- PerFRDiff: Personalised Weight Editing for Multiple Appropriate Facial Reaction GenerationHengde Zhu, Xiangyu Kong, Weicheng Xie, Xin Huang et al.ACM MM 2024 · 12 citations
- ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion ModelCheng Luo, Siyang Song, Siyuan Yan, Zhen Yu et al.ACM MM 2025 · 1 citation
- PerReactor: Offline Personalised Multiple Appropriate Facial Reaction GenerationHengde Zhu, Xiangyu Kong, Weicheng Xie, Xin Huang et al.AAAI 2025 · 3 citations
- PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic InteractionZhi-Yi Lin, Thomas Markhorst, Jouh Yeong Chew, Xucong ZhangCVPR 2026 · 3 citations
- SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationWenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang et al.CVPR 2023
