WaveEx: Accelerating Flow Matching-based Speech Generation via Wavelet-guided Extrapolation
Xiaoqian Liu, Xiyan Gui, Zhengkun Ge, Yuan Ge, Chang Zou, Jiacheng Liu, Zhikang Niu, Qixi Zheng, Chen Xu, Xie Chen, Tong Xiao, Jingbo Zhu, Linfeng Zhang
摘要
Flow matching-based generative models offer a principled approach to modeling continuous-time dynamics in speech generation. However, inference is often computationally expensive due to repeated neural network evaluations required by ODE solvers. We propose WaveEx, a training-free and plug-in acceleration framework which replaces portions of ODE integration with wavelet-guided extrapolation. By leveraging the multi-scale structure of latent trajectories, WaveEx predicts future states directly in the frequency domain without additional model evaluations or architectural changes. WaveEx consistently accelerates inference across diverse speech generation tasks. The gains are especially pronounced in tasks like speech synthesis (up to 5.73× speedup) and music generation (2.75×), where flow matching plays a central role in alignment modeling and dense ODE integration. Even in tasks with simpler input-output mappings such as speech enhancement (4.55×) and voice conversion (2.75×), WaveEx still achieves notable acceleration, demonstrating the robustness and generalizability of the approach. These results highlight wavelet-guided extrapolation as a lightweight and broadly applicable alternative to full ODE solving for flow matching-based speech generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Improved Distribution Matching Distillation for Fast Image SynthesisTianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang 等NeurIPS 2024 · 被引用 728 次
- Voicebox: Text-Guided Multilingual Universal Speech Generation at ScaleMatthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer 等NeurIPS 2023 · 被引用 613 次
- NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing SynthesizersKai Shen, Zeqian Ju, Xu Tan, Eric Liu 等ICLR 2024 · 被引用 362 次
相关 Paper
- FlowCast: Trajectory Forecasting for Scalable Zero-Cost Speculative Flow MatchingDivya Jyoti Bajpai, Shubham Agarwal, Apoorv Saxena, Kuldeep Kulkarni 等ICLR 2026 · 被引用 3 次
- FastFlow: Accelerating The Generative Flow Matching Models with Bandit InferenceDivya Jyoti Bajpai, Dhruv Bhardwaj, Soumya Roy, Tejas Duseja 等ICLR 2026 · 被引用 3 次
- Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion ModelXiangyu Zhang, Daijiao Liu, Hexin Liu, Qiquan Zhang 等EMNLP 2024 · 被引用 3 次
- A-FloPS: Accelerating Diffusion Models via Adaptive Flow Path SamplerCheng Jin, Zhenyu Xiao, Yuantao GuAAAI 2026
- Bi-Anchor Interpolation Solver for Accelerating Generative ModelingHongxu CHEN, Hongxiang Li, Zhen Wang, Long ChenICML 2026 · 被引用 3 次
