Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation
Dingcheng Zhen, Shunshun Yin, Shiyang Qin, Hou Yi, Ziwei Zhang, Siyuan Liu, Gan Qi, Ming Tao
2025Year
3Top-tier citations
Abstract
Stage 2 Efficient Temporal Module Before Temporal Module After Temporal Module Real-Time Performance 200ms chunks 🚀 For 200ms Audio Chunk Condition Total Inference Time < 200ms 🚀 Audio Encode cost 7ms Stage2 cost 71ms Stage1 cost 106ms Diverse Move Figure 1. Teller framework is the first autoregressive framework for real-time, audio-driven portrait animation, achieving up to 25 FPS while preserving realistic body part and accessory movements. Demo can be found at https://teller-avatar.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generationxiang deng, Feng Gao, Yong Zhang, Youxin Pang et al.CVPR 2026 · 2 citations
- EchoAvatar: Real-time Generative Avatar Animation from Audio StreamsBohong Chen, Yumeng Li, Yinglin Xu, Youyi Zheng et al.SIGGRAPH 2026
- LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video CreationWenhui Song, Hanhui Li, Jiehui Huang, Panwen Hu et al.ACM MM 2025
Builds on21
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Any-to-Any Generation via Composable DiffusionZineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng et al.NeurIPS 2023 · 294 citations
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang et al.NeurIPS 2024 · 253 citations
- EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark ConditionsZhiyuan Chen, Jiajiong Cao, Zhiquan Chen, Yuming Li et al.AAAI 2025 · 197 citations
Related papers
- READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head GenerationHaotian Wang, Yuzhe Weng, Jun Du, Haoran Xu et al.AAAI 2026 · 1 citation
- REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming DistillationHaotian Wang, Yuzhe Weng, Jun Du, Haoran Xu et al.ICML 2026 · 4 citations
- TACR-Net: Editing on Deep Video and Voice PortraitsLuchuan Song, Bin Liu, Guojun Yin, Xiaoyi Dong et al.ACM MM 2021 · 20 citations
- INFP: Audio-Driven Interactive Head Generation in Dyadic ConversationsYongming Zhu, Longhao Zhang, Zhengkun Rong, Tianshu Hu et al.CVPR 2025
- StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human AvatarsZhiyao Sun, Ziqiao Peng, Yifeng Ma, Yi Chen et al.CVPR 2026 · 26 citations
