A3T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and Editing
He Bai, Renjie Zheng, Jun-Kun Chen, Mingbo Ma, Xintong Li, Liang Huang
Abstract
Recently, speech representation learning has improved many speech-related tasks such as speech recognition, speech classification, and speech-to-text translation. However, all the above tasks are in the direction of speech understanding, but for the inverse direction, speech synthesis, the potential of representation learning is yet to be realized, due to the challenging nature of generating high-quality speech. To address this problem, we propose our framework, Alignment-Aware Acoustic-Text Pretraining (AT), which reconstructs masked acoustic signals with text input and acoustic-text alignment during training. In this way, the pretrained model can generate high quality reconstructed spectrogram, which can be applied to the speech editing and unseen speaker TTS directly. Experiments show AT outperforms SOTA models on speech editing, and improves multi-speaker speech synthesis without the external speaker verification model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4064c4a-acff-4699-995b-2ce1c8144550Cited by top-tier papers10
- Voicebox: Text-Guided Multilingual Universal Speech Generation at ScaleMatthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer et al.NeurIPS 2023 · 613 citations
- Proactive Detection of Voice Cloning with Localized WatermarkingRobin San Roman, Pierre Fernandez, Hady Elsahar, Alexandre Défossez et al.ICML 2024 · 119 citations
- P-Flow: A Fast and Data-Efficient Zero-Shot TTS through Speech PromptingSungwon Kim, Kevin J. Shih, Rohan Badlani, João Felipe Santos et al.NeurIPS 2023 · 75 citations
- Generative Pre-training for Speech with Flow MatchingAlexander H. Liu, Matthew Le, Apoorv Vyas, Bowen Shi et al.ICLR 2024 · 66 citations
- SongCreator: Lyrics-based Universal Song GenerationShun Lei, Yixuan Zhou, Boshi Tang, Max W. Y. Lam et al.NeurIPS 2024 · 33 citations
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech TranslationRenjie Zheng, Jun-Kun Chen, Mingbo Ma, Liang HuangICML 2021 · 74 citations
- Segatron: Segment-Aware Transformer for Language Modeling and UnderstandingHe Bai, Peng Shi, Jimmy Lin, Yuqing Xie et al.AAAI 2021 · 25 citations
Related papers
- SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language ProcessingJunyi Ao, Rui Wang, Long Zhou, Chengyi Wang et al.ACL 2022
- UniWav: Towards Unified Pre-training for Speech Representation Learning and GenerationAlexander H. Liu, Sang-gil Lee, Chao-Han Huck Yang, Yuan Gong et al.ICLR 2025
- SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-trainingZiqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu et al.EMNLP 2022 · 38 citations
- CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-TrainingZhenhui Ye, Rongjie Huang, Yi Ren, Ziyue Jiang et al.ACL 2023 · 13 citations
- UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled DataChengyi Wang, Yu Wu, Yao Qian, Ken'ichi Kumatani et al.ICML 2021 · 140 citations
