REA-Listener: Real-Time Listening Head Generation with Dynamic Emotion Modeling and Flexible Modality Adaptation
Sizhe Zhao, Chenyang Wang, Weiyu Zhao, Zonglin Li, Ming Li, Shengping Zhang
Abstract
Listening head generation aims to synthesize realistic and responsive non-verbal listener head motions that respond to speakers in conversational scenarios. Existing methods typically rely on fixed audio-visual input modalities and predefined emotion labels, limiting their adaptability and expressiveness in real-world scenarios. In this paper, we propose a novel real-time framework, REA-Listener, to generate high-fidelity listening head videos with flexible modality adaptation and dynamic emotion modeling. Specifically, we first propose a Modality-Adaptive Mixture of Experts (MA-MoE) module to encode arbitrary combinations of speaker audio and visual signals into a unified embedding space, ensuring robustness under partial modality conditions. To further enhance the temporal consistency of listener emotion, we present a lightweight emotional head dynamics generator with a multi-modal emotion predictor, which infers listener emotions dynamically from speaker context alongside head motion coefficient prediction. Finally, we employ a 3D-aware renderer based on 3D Gaussian Splatting to produce high-quality listener head videos in real time. With these components, our approach achieves efficient head motion generation at 30fps on a single NVIDIA RTX 3090 GPU, supporting real-time interaction. Extensive evaluations and applications demonstrate that our method outperforms state-of-the-art methods in listening head generation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive InteractionShiying Li, Xingqun Qi, Bingkun Yang, Weile Chen et al.AAAI 2026 · 2 citations
- CustomListener: Text-Guided Responsive Interaction for User-Friendly Listening Head GenerationXi Liu, Ying Guo, Cheng Zhen, Tong Li et al.CVPR 2024
- Emotional Listener Portrait: Realistic Listener Motion Simulation in ConversationLuchuan Song, Guojun Yin, Zhenchao Jin, Xiaoyi Dong et al.ICCV 2023 · 19 citations
- MFR-Net: Multi-faceted Responsive Listening Head Generation via Denoising Diffusion ModelJin Liu, Xi Wang, Xiaomeng Fu, Yesheng Chai et al.ACM MM 2023 · 10 citations
- Diffusion-based Realistic Listening Head Generation via Hybrid Motion ModelingYinuo Wang, Yanbo Fan, Xuan Wang, Yu Guo et al.CVPR 2025
