CustomListener: Text-Guided Responsive Interaction for User-Friendly Listening Head Generation
Xi Liu, Ying Guo, Cheng Zhen, Tong Li, Yingying Ao, Pengfei Yan
摘要
Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion. The applications of listener agent generation in virtual interaction have promoted many works achieving diverse and fine-grained motion generation. However, they can only manipulate motions through simple emotional labels, but cannot freely control the listener's motions. Since listener agents should have human-like attributes (e.g. identity, personality) which can be freely customized by users, this limits their realism. In this paper, we propose a user-friendly framework called CustomListener to realize the free-form text prior guided listener generation. To achieve speaker-listener coordination, we design a Static to Dynamic Portrait module (SDP), which interacts with speaker information to transform static text into dynamic portrait token with completion rhythm and amplitude information. To achieve coherence between segments, we design a Past Guided Generation module (PGG) to maintain the consistency of customized listener attributes through the motion prior, and utilize a diffusion-based structure conditioned on the portrait token and the motion prior to realize the controllable generation. To train and evaluate our model, we have constructed two text-annotated listening head datasets based on ViCo and RealTalk, which provide text-video paired labels. Extensive experiments have verified the effectiveness of our model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human AvatarsZhiyao Sun, Ziqiao Peng, Yifeng Ma, Yi Chen 等CVPR 2026 · 被引用 26 次
- Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural ConversationTaekyung Ki, Sangwon Jang, Jaehyeong Jo, Jaehong Yoon 等CVPR 2026 · 被引用 18 次
- UniLS: End-to-End Audio-Driven Avatars for Unified Listening and SpeakingXuangeng Chu, Ruicong Liu, Yifei Huang, Yun Liu 等CVPR 2026 · 被引用 12 次
- VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive InteractionShiying Li, Xingqun Qi, Bingkun Yang, Weile Chen 等AAAI 2026 · 被引用 2 次
- Ditailistener: Controllable High Fidelity Listener Video Generation with DiffusionMaksim Siniukov, Di Chang, Minh Tran, Hongkun Gong 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper15
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 被引用 701 次
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 被引用 672 次
- PIRenderer: Controllable Portrait Image Generation via Semantic Neural RenderingYurui Ren, Ge Li, Yuanqi Chen, Thomas H. Li 等ICCV 2021 · 被引用 284 次
相关 Paper
- Emotional Listener Portrait: Realistic Listener Motion Simulation in ConversationLuchuan Song, Guojun Yin, Zhenchao Jin, Xiaoyi Dong 等ICCV 2023 · 被引用 19 次
- REA-Listener: Real-Time Listening Head Generation with Dynamic Emotion Modeling and Flexible Modality AdaptationSizhe Zhao, Chenyang Wang, Weiyu Zhao, Zonglin Li 等ACM MM 2025
- Diffusion-based Realistic Listening Head Generation via Hybrid Motion ModelingYinuo Wang, Yanbo Fan, Xuan Wang, Yu Guo 等CVPR 2025
- MFR-Net: Multi-faceted Responsive Listening Head Generation via Denoising Diffusion ModelJin Liu, Xi Wang, Xiaomeng Fu, Yesheng Chai 等ACM MM 2023 · 被引用 10 次
- ARIG: Autoregressive Interactive Head Generation for Real-Time ConversationsYing Guo, Xi Liu, Cheng Zhen, Pengfei Yan 等ICCV 2025
