VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
Puyuan Peng, Po-Yao Huang, Shang-Wen Li, Abdelrahman Mohamed, David Harwath
摘要
We introduce VOICECRAFT, a token infilling neural codec language model, that achieves state-of-the-art performance on both speech editing and zero-shot text-to-speech (TTS) on audiobooks, internet videos, and podcasts 1 . VOICECRAFT employs a Transformer decoder architecture and introduces a token rearrangement procedure that combines causal masking and delayed stacking to enable generation within an existing sequence. On speech editing tasks, VOICECRAFT produces edited speech that is nearly indistinguishable from unedited recordings in terms of naturalness, as evaluated by humans; for zero-shot TTS, our model outperforms prior SotA models including VALL-E and the popular commercial model XTTS v2. Crucially, the models are evaluated on challenging and realistic datasets, that consist of diverse accents, speaking styles, recording conditions, and background noise and music, and our model performs consistently well compared to other models and real recordings. In particular, for speech editing evaluation, we introduce a high quality, challenging, and realistic dataset named REALEDIT. We encourage readers to listen to the demos at https: //jasonppy.github.io/VoiceCraft_web .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper38
- Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language ModelZhen Ye, Peiwen Sun, Jiahe Lei, Hongzhan Lin 等AAAI 2025 · 被引用 89 次
- UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal InteractionsGuozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng 等CVPR 2026 · 被引用 40 次
- SongCreator: Lyrics-based Universal Song GenerationShun Lei, Yixuan Zhou, Boshi Tang, Max W. Y. Lam 等NeurIPS 2024 · 被引用 33 次
- Metis: A Foundation Speech Generation Model with Masked Generative Pre-trainingYuancheng Wang, Jiachen Zheng, Junan Zhang, Xueyao Zhang 等NeurIPS 2025 · 被引用 25 次
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language ModelingYuancheng Wang, Dekun Chen, Xueyao Zhang, Junan Zhang 等NeurIPS 2025 · 被引用 22 次
它引用的顶会 Paper14
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 被引用 1,267 次
- Voicebox: Text-Guided Multilingual Universal Speech Generation at ScaleMatthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer 等NeurIPS 2023 · 被引用 613 次
相关 Paper
- VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech EditingZhisheng Zheng, Puyuan Peng, Anuj Diwan, Cong Phuoc Huynh 等EMNLP 2025
- MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec TransformerYuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng 等ICLR 2025
- UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and VocodingChenpeng Du, Yiwei Guo, Feiyu Shen, Zhijun Liu 等AAAI 2024 · 被引用 64 次
- Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech SynthesisYifan Yang, Shujie Liu, Jinyu Li, Yuxuan Hu 等ACM MM 2025 · 被引用 1 次
- ELLA-V: Stable Neural Codec Language Modeling with Alignment-Guided Sequence ReorderingYakun Song, Zhuo Chen, Xiaofei Wang, Ziyang Ma 等AAAI 2025 · 被引用 75 次
