ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion Model
Cheng Luo, Siyang Song, Siyuan Yan, Zhen Yu, Zongyuan Ge
Abstract
The automatic generation of diverse and human-like facial reactions in dyadic dialogue remains a critical challenge for human-computer interaction systems. Existing methods fail to model the stochasticity and dynamics inherent in real human reactions. To address this, we propose ReactDiff, a novel temporal diffusion framework for generating diverse facial reactions that are appropriate for responding to any given dialogue context. Our key insight is that plausible human reactions demonstrate smoothness, and coherence over time, and conform to constraints imposed by human facial anatomy. To achieve this, ReactDiff incorporates two vital priors (spatio-temporal facial kinematics) into the diffusion process: i) temporal facial behavioral kinematics and ii) facial action unit dependencies. These two constraints guide the model toward realistic human reaction manifolds, avoiding visually unrealistic jitters, unstable transitions, unnatural expressions, and other artifacts. Extensive experiments on the REACT2024 dataset demonstrate that our approach not only achieves state-of-the-art reaction quality but also excels in diversity and reaction appropriateness. Our code is publicly available at https://github.com/lingjivoo/ReactDiff.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b06aa0c9-33f0-422d-a40b-9d1a647dd67dCited by top-tier papers2
- DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture GenerationYICHEN PENG, Jyun-Ting Song, Siyeol Jung, RUOFAN LIU et al.CVPR 2026 · 7 citations
- OmniResponse: Online Multimodal Conversational Response Generation in Dyadic InteractionsCheng Luo, Jianghui Wang, Bing Li, Siyang Song et al.NeurIPS 2025 · 4 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
Related papers
- Smooth Online Multiple Appropriate Facial Reaction GenerationWeicheng Xie, Chunlin Yan, Siyang Song, Zitong Yu et al.ACM MM 2025
- ReGenNet: Towards Human Action-Reaction SynthesisLiang Xu, Yizhou Zhou, Yichao Yan, Xin Jin et al.CVPR 2024 · 18 citations
- MAUGen: A Unified Diffusion Approach for Multi-Identity Facial Expression and AU Label GenerationXiangdong Li, Ye Lou, Ao Gao, Wei Zhang et al.AAAI 2026
- From Audio to Photoreal Embodiment: Synthesizing Humans in ConversationsEvonne Ng, Javier Romero, Timur M. Bagautdinov, Shaojie Bai et al.CVPR 2024 · 37 citations
- PerReactor: Offline Personalised Multiple Appropriate Facial Reaction GenerationHengde Zhu, Xiangyu Kong, Weicheng Xie, Xin Huang et al.AAAI 2025 · 3 citations
