ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models
Sang-Hoon Lee, Ha-Yeong Choi
摘要
Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit generative capacity. To address this issue, we propose ReGen, a hierarchical multi-prompt representation generation framework that jointly estimates multiple vector fields for both representations and data within a single diffusion model. We further introduce generalized flow matching (GFM) to improve the generalization of conditional flow matching (CFM). We validate ReGen on single-stage waveform diffusion models including neural audio codec and Wave-VAE. ReGen significantly improves waveform generation quality from highly compressed latent representations at 12.5 Hz. We also present ReGenVoice, a latent diffusion model (LDM)based text-to-speech model that achieves strong speech intelligibility (WER) and speaker similarity (SIM) with a small dataset. Moreover, operating the LDM at 6.25 Hz with rich semantic and acoustic latent representation enables efficient training and sampling, requiring only 1 day of training on 4 GPUs and fast inference with an RTF of 0.08. Audio samples are available at https: //regenvoice.github.io/demo/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- Diffusion Transformers with Representation AutoencodersBoyang Zheng, Nanye Ma, Shengbang Tong, Saining XieICLR 2026 · 被引用 288 次
相关 Paper
- PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform GenerationSang-Hoon Lee, Ha-Yeong Choi, Seong-Whan LeeICLR 2025
- Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You ThinkGe Wu, Shen Zhang, Ruijing Shi, Shanghua Gao 等NeurIPS 2025 · 被引用 102 次
- SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech SynthesisYifan Liang, Andong Li, Kang Yang, Guochen Yu 等AAAI 2026
- Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow MatchingYongqi Wang, Wenxiang Guo, Rongjie Huang, Jiawei Huang 等NeurIPS 2024 · 被引用 73 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
