StreamDiffusion: A Pipeline-Level Solution for Real-Time Interactive Generation
Akio Kodaira, Chenfeng Xu, Toshiki Hazama, Takanori Yoshimoto, Kohei Ohno, Shogo Mitsuhori, Soichi Sugano, Hanying Cho, Zhijian Liu, Masayoshi Tomizuka, Kurt Keutzer
Abstract
We introduce StreamDiffusion, a real-time diffusion pipeline designed for interactive image generation. Existing diffusion models are adept at creating images from text or image prompts, yet they often fall short in real-time interaction. This limitation becomes particularly evident in scenarios involving continuous input, such as Metaverse, live video streaming, and broadcasting, where high throughput is imperative. To address this, we present a novel approach that transforms the original sequential denoising into the batching denoising process. Stream Batch eliminates the conventional wait-and-interact approach and enables fluid and high throughput streams. To handle the frequency disparity between data input and model throughput, we design a novel input-output queue for parallelizing the streaming process. Moreover, the existing diffusion pipeline uses classifier-free guidance(CFG), which requires additional U-Net computation. To mitigate the redundant computations, we propose a novel residual classifier-free guidance (RCFG) algorithm that reduces the number of negative conditional denoising steps to only one or even zero. Besides, we introduce a stochastic similarity filter(SSF) to optimize power consumption. Our Stream Batch achieves around 1.5x speedup compared to the sequential denoising method at different denoising levels. The proposed RCFG leads to speeds up to 2.05x higher than the conventional CFG. Combining the proposed strategies and existing mature acceleration tools makes the image-to-image generation achieve up-to 91.07fps on one RTX4090, improving the throughputs of AutoPipline developed by Diffusers over 59.56x. Furthermore, our proposed StreamDiffusion also significantly reduces the energy consumption by 2.39x on one RTX3060 and 1.99x on one RTX4090, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- The Matrix: Infinite-Horizon World Generation with Real-Time Moving ControlRuili Feng, Han Zhang, Zhilei Shu, Zhantao Yang et al.NeurIPS 2025 · 92 citations
- Autoregressive Adversarial Post-Training for Real-Time Interactive Video GenerationShanchuan Lin, Ceyuan Yang, Hao He, Jianwen Jiang et al.NeurIPS 2025 · 89 citations
- StreamDiT: Real-Time Streaming Text-to-Video GenerationAkio Kodaira, Tingbo Hou, Ji Hou, Markos Georgopoulos et al.CVPR 2026 · 45 citations
- Media2Face: Co-speech Facial Animation Generation With Multi-Modality GuidanceQingcheng Zhao, Pengyu Long, Qixuan Zhang, Dafei Qin et al.SIGGRAPH 2024 · 40 citations
- WorldLens: Full-Spectrum Evaluations of Driving World Models in Real WorldAo Liang, Lingdong Kong, Tianyi Yan, Hongsi Liu et al.CVPR 2026 · 28 citations
Builds on23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- FasterCache: Training-Free Video Diffusion Model Acceleration with High QualityZhengyao Lv, Chenyang Si, Junhao Song, Zhenyu Yang et al.ICLR 2025
- No Training, No Problem: Rethinking Classifier-Free Guidance for Diffusion ModelsSeyedmorteza Sadat, Manuel Kansy, Otmar Hilliges, Romann M. WeberICLR 2025
- SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion ModelsJaerin Lee, Daniel Sungho Jung, Kanggeon Lee, Kyoung Mu LeeCVPR 2025
- FAST-AR: Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse AttentionDvir Samuel, Issar Tzachor, Matan Levy, Michael Green et al.ICML 2026 · 7 citations
- Rethinking the Spatial Inconsistency in Classifier-Free Diffusion GuidanceDazhong Shen, Guanglu Song, Zeyue Xue, Fu-Yun Wang et al.CVPR 2024 · 12 citations
