StreamFlow: Streaming Audio Generation from Discrete Tokens via Streaming Flow Matching
Ha-Yeong Choi, Sang-Hoon Lee
摘要
Diffusion models have demonstrated remarkable generative capabilities, and Conditional Flow Matching (CFM) has improved their inference efficiency by following optimal transport paths. However, CFM-based models still require multiple iterative sampling steps, which makes them unsuitable for real-time or streaming generation scenarios. In this paper, we introduce StreamFlow, a novel streaming generative model designed for real-time audio generation from discrete tokens. StreamFlow leverages a causal noising training framework along the time axis and predicts multi-time vector fields at once on each stream, enabling streaming inference with minimal latency. To further improve generalization, we propose Scale-DiT, a Diffusion Transformer architecture that enhances robustness by modeling, normalizing, and scaling feature differences prior to skip connections. This significantly improves the robustness and performance of DiT without increasing the parameter size. We validate the effectiveness of StreamFlow through audio reconstruction tasks using discrete tokens from EnCodec and Mimi, demonstrating both high-fidelity synthesis and streaming capability. Furthermore, we successfully incorporated our model into fully-duplex streaming speech language models of Moshi by replacing the Mimi decoder. Audio samples are available at https://streamflow25.github.io/demo/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
相关 Paper
- StreamDiT: Real-Time Streaming Text-to-Video GenerationAkio Kodaira, Tingbo Hou, Ji Hou, Markos Georgopoulos 等CVPR 2026 · 被引用 45 次
- READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head GenerationHaotian Wang, Yuzhe Weng, Jun Du, Haoran Xu 等AAAI 2026 · 被引用 1 次
- Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head SynthesisTianqi Li, Ruobing Zheng, Minghui Yang, Jingdong Chen 等ACM MM 2025 · 被引用 4 次
- Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEsSicheng Xu, Yu Deng, Shoukang Hu, Yichuan Wang 等CVPR 2026 · 被引用 1 次
- ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion ModelsSang-Hoon Lee, Ha-Yeong ChoiICML 2026
