CustomListener: Text-Guided Responsive Interaction for User-Friendly Listening Head Generation
Xi Liu, Ying Guo, Cheng Zhen, Tong Li, Yingying Ao, Pengfei Yan
Abstract
Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion. The applications of listener agent generation in virtual interaction have promoted many works achieving diverse and fine-grained motion generation. However, they can only manipulate motions through simple emotional labels, but cannot freely control the listener's motions. Since listener agents should have human-like attributes (e.g. identity, personality) which can be freely customized by users, this limits their realism. In this paper, we propose a user-friendly framework called CustomListener to realize the free-form text prior guided listener generation. To achieve speaker-listener coordination, we design a Static to Dynamic Portrait module (SDP), which interacts with speaker information to transform static text into dynamic portrait token with completion rhythm and amplitude information. To achieve coherence between segments, we design a Past Guided Generation module (PGG) to maintain the consistency of customized listener attributes through the motion prior, and utilize a diffusion-based structure conditioned on the portrait token and the motion prior to realize the controllable generation. To train and evaluate our model, we have constructed two text-annotated listening head datasets based on ViCo and RealTalk, which provide text-video paired labels. Extensive experiments have verified the effectiveness of our model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d774b2c1-222d-4ee4-8b77-5b59826c1bd8Cited by top-tier papers11
- StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human AvatarsZhiyao Sun, Ziqiao Peng, Yifeng Ma, Yi Chen et al.CVPR 2026 · 26 citations
- Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural ConversationTaekyung Ki, Sangwon Jang, Jaehyeong Jo, Jaehong Yoon et al.CVPR 2026 · 18 citations
- UniLS: End-to-End Audio-Driven Avatars for Unified Listening and SpeakingXuangeng Chu, Ruicong Liu, Yifei Huang, Yun Liu et al.CVPR 2026 · 12 citations
- VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive InteractionShiying Li, Xingqun Qi, Bingkun Yang, Weile Chen et al.AAAI 2026 · 2 citations
- Ditailistener: Controllable High Fidelity Listener Video Generation with DiffusionMaksim Siniukov, Di Chang, Minh Tran, Hongkun Gong et al.ICCV 2025 · 1 citation
Builds on15
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
- PIRenderer: Controllable Portrait Image Generation via Semantic Neural RenderingYurui Ren, Ge Li, Yuanqi Chen, Thomas H. Li et al.ICCV 2021 · 284 citations
Related papers
- Emotional Listener Portrait: Realistic Listener Motion Simulation in ConversationLuchuan Song, Guojun Yin, Zhenchao Jin, Xiaoyi Dong et al.ICCV 2023 · 19 citations
- REA-Listener: Real-Time Listening Head Generation with Dynamic Emotion Modeling and Flexible Modality AdaptationSizhe Zhao, Chenyang Wang, Weiyu Zhao, Zonglin Li et al.ACM MM 2025
- Diffusion-based Realistic Listening Head Generation via Hybrid Motion ModelingYinuo Wang, Yanbo Fan, Xuan Wang, Yu Guo et al.CVPR 2025
- MFR-Net: Multi-faceted Responsive Listening Head Generation via Denoising Diffusion ModelJin Liu, Xi Wang, Xiaomeng Fu, Yesheng Chai et al.ACM MM 2023 · 10 citations
- ARIG: Autoregressive Interactive Head Generation for Real-Time ConversationsYing Guo, Xi Liu, Cheng Zhen, Pengfei Yan et al.ICCV 2025
