UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
Alexander H. Liu, Sang-gil Lee, Chao-Han Huck Yang, Yuan Gong, Yu-Chiang Frank Wang, James R. Glass, Rafael Valle, Bryan Catanzaro
摘要
Pre-training and representation learning have been playing an increasingly important role in modern speech processing. Nevertheless, different applications have been relying on different foundation models, since predominant pre-training techniques are either designed for discriminative tasks or generative tasks. In this work, we make the first attempt at building a unified pre-training framework for both types of tasks in speech. We show that with the appropriate design choices for pre-training, one can jointly learn a representation encoder and generative audio decoder that can be applied to both types of tasks. We propose UniWav, an encoder-decoder framework designed to unify pre-training representation learning and generative tasks. On speech recognition, text-to-speech, and speech tokenization, UniWav achieves comparable performance to different existing foundation models, each trained on a specific task. Our findings suggest that a single generalpurpose foundation model for speech can be built to replace different foundation models, reducing the overhead and cost of pre-training. Audio demo page: https://alexander-h-liu.github.io/uniwav-demo.github.io/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language ModelingLiang-Hsuan Tseng, Yi-Chang Chen, Kuan Yi Lee, Da-shan Shiu 等ICLR 2026 · 被引用 26 次
- AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion ForcingWilliam Chen, Prem Seetharaman, Rithesh Kumar, Oriol Nieto 等ICML 2026 · 被引用 7 次
- From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint TrainingTianqiao Liu, Xueyi Li, Hao Wang, Haoxuan Li 等ICLR 2026 · 被引用 6 次
- Alethia: a Foundational Encoder for Voice DeepfakesYi Zhu, Brahmi Dwivedi, Jayaram Raghuram, Surya KoppisettiICML 2026
- PACE: Pretrained Audio Continual LearningChang Li, Kanglei Zhou, Liyuan WangICLR 2026
它引用的顶会 Paper23
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
相关 Paper
- SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language ProcessingJunyi Ao, Rui Wang, Long Zhou, Chengyi Wang 等ACL 2022
- A3T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and EditingHe Bai, Renjie Zheng, Jun-Kun Chen, Mingbo Ma 等ICML 2022 · 被引用 64 次
- Generative Pre-training for Speech with Flow MatchingAlexander H. Liu, Matthew Le, Apoorv Vyas, Bowen Shi 等ICLR 2024 · 被引用 66 次
- UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled DataChengyi Wang, Yu Wu, Yao Qian, Ken'ichi Kumatani 等ICML 2021 · 被引用 140 次
- UniAudio: Towards Universal Audio Generation with Large Language ModelsDongchao Yang, Jinchuan Tian, Xu Tan, Rongjie Huang 等ICML 2024 · 被引用 54 次
