StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars
Zhiyao Sun, Ziqiao Peng, Yifeng Ma, Yi Chen, zhengguang zhou, Zixiang Zhou, Guozhen Zhang, Youliang Zhang, Yuan Zhou, qinglin lu, Yong-Jin Liu
摘要
Real-time, streaming interactive avatars represent a critical yet challenging goal in digital human research. Although diffusion-based human avatar generation methods achieve remarkable success, their non-causal architecture and high computational costs make them unsuitable for streaming. Moreover, existing interactive approaches are typically limited to head-and-shoulder region, limiting their ability to produce gestures and body motions. To address these challenges, we propose a two-stage autoregressive adaptation and acceleration framework that applies autoregressive distillation and adversarial refinement to adapt a high-fidelity human video diffusion model for real-time, interactive streaming. To ensure long-term stability and consistency, we introduce three key components: a Reference Sink, a Reference-Anchored Positional Re-encoding (RAPR) strategy, and a Consistency-Aware Discriminator. Building on this framework, we develop a one-shot, interactive, human avatar model capable of generating both natural talking and listening behaviors with coherent gestures. Extensive experiments demonstrate that our method achieves state-of-the-art performance, surpassing existing approaches in generation quality, real-time efficiency, and interaction naturalness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- OmniSync: Towards Universal Lip Synchronization via Diffusion TransformersZiqiao Peng, Jiwen Liu, Haoxian Zhang, Xiaoqiang Liu 等NeurIPS 2025 · 被引用 30 次
- Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video GenerationHongzhou Zhu, Min Zhao, Guande He, Hang Su 等ICML 2026
它引用的顶会 Paper38
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz 等NeurIPS 2024 · 被引用 751 次
相关 Paper
- REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming DistillationHaotian Wang, Yuzhe Weng, Jun Du, Haoran Xu 等ICML 2026 · 被引用 4 次
- ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal GuidanceHaijie Yang, Zhenyu Zhang, Hao Tang, Jianjun Qian 等ACM MM 2024 · 被引用 3 次
- Expressive Talking Human from Single-Image with Imperfect PriorsJun Xiang, Yudong Guo, Leipeng Hu, Boyang Guo 等ICCV 2025 · 被引用 3 次
- Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEsSicheng Xu, Yu Deng, Shoukang Hu, Yichuan Wang 等CVPR 2026 · 被引用 1 次
- InstantViR: Real-Time Video Inverse Problem Solver with Distilled Diffusion PriorWeimin Bai, Suzhe Xu, Yiwei Ren, Jinhua Hao 等CVPR 2026 · 被引用 3 次
