Scaling Rich Style-Prompted Text-to-Speech Datasets
Anuj Diwan, Zhisheng Zheng, David Harwath, Eunsol Choi
摘要
We introduce Paralinguistic Speech Captions (ParaSpeechCaps), a large-scale dataset that annotates speech utterances with rich style captions. While rich abstract tags (e.g. guttural, nasal, pained) have been explored in small-scale human-annotated datasets, existing large-scale datasets only cover basic tags (e.g. low-pitched, slow, loud). We combine off-the-shelf text and speech embedders, classifiers and an audio language model to automatically scale rich tag annotations for the first time. ParaSpeechCaps covers a total of 59 style tags, including both speaker-level intrinsic tags and utterance-level situational tags. It consists of 282 hours of human-labelled data (PSC-Base) and 2427 hours of automatically annotated data (PSC-Scaled). We finetune Parler-TTS, an open-source style-prompted TTS model, on ParaSpeechCaps, and achieve improved style consistency (+7.9% Consistency MOS) and speech quality (+15.5% Naturalness MOS) over the best performing baseline that combines existing rich style tag datasets. We ablate several of our dataset design choices to lay the foundation for future work in this space. Our dataset, models and code are released at https://github. com/ajd12342/paraspeechcaps .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language InstructionsDekun Chen, Xueyao Zhang, Yuancheng Wang, Kenan Dai 等ICLR 2026 · 被引用 21 次
- Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-trainingYifan Yang, Bing Han, Hui Wang, Wei Wang 等ACL 2026 · 被引用 4 次
- Revisiting Audio-language Pretraining for Learning General-purpose Audio RepresentationWei-Cheng Tseng, Xuanru Zhou, Mingyue Huo, Yiwen Shao 等ACL 2026 · 被引用 2 次
- Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training MethodsXuanru Zhou, Yiwen Shao, Wei-Cheng Tseng, Dong YuCVPR 2026 · 被引用 1 次
它引用的顶会 Paper6
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- PromptTTS 2: Describing and Generating Voices with Text PromptYichong Leng, Zhifang Guo, Kai Shen, Zeqian Ju 等ICLR 2024 · 被引用 80 次
- MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech SynthesisWenhao Guan, Yishuang Li, Tao Li, Hukai Huang 等AAAI 2024 · 被引用 25 次
- SpeechCraft: A Fine-Grained Expressive Speech Dataset with Natural Language DescriptionZeyu Jin, Jia Jia, Qixin Wang, Kehan Li 等ACM MM 2024 · 被引用 12 次
相关 Paper
- DNASpeech: A Contextualized and Situated Text-to-Speech Dataset with Dialogues, Narratives and ActionsChuanqi Cheng, Hongda Sun, Bo Du, Shuo Shang 等ACL 2025
- Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio CaptioningSangyeon Cho, Mingi Kim, Jinkwon Hwang, Jaehoon Go 等EMNLP 2025
- UniSS: Unified Expressive Speech-to-Speech Translation with Your VoiceSitong Cheng, Bianweizhen, Xinsheng Wang, Ruibin Yuan 等ICLR 2026 · 被引用 7 次
- ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech InteractionShu-Wen Yang, Ming Tu, Ting-Wei Liu, Xinghua Qu 等ICLR 2026 · 被引用 29 次
- QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and DescriptionsSiyin Wang, Wenyi Yu, Xianzhao Chen, Xiaohai Tian 等ACL 2025 · 被引用 20 次
