DNASpeech: A Contextualized and Situated Text-to-Speech Dataset with Dialogues, Narratives and Actions
Chuanqi Cheng, Hongda Sun, Bo Du, Shuo Shang, Xinrong Hu, Rui Yan
Abstract
In this paper, we propose contextualized and situated text-to-speech (CS-TTS), a novel TTS task to promote more accurate and customized speech generation using prompts with Dialogues, Narratives, and Actions (DNA). While prompt-based TTS methods facilitate controllable speech generation, existing TTS datasets lack situated descriptive prompts aligned with speech data. To address this data scarcity, we develop an automatic annotation pipeline enabling multifaceted alignment among speech clips, content text, and their respective descriptions. Based on this pipeline, we present DNASpeech, a novel CS-TTS dataset with high-quality speeches with DNA prompt annotations. DNASpeech contains 2,395 distinct characters, 4,452 scenes, and 22,975 dialogue utterances, along with over 18 hours of high-quality speech recordings. To accommodate more specific task scenarios, we establish a leaderboard featuring two new subtasks for evaluation: CS-TTS with narratives and CS-TTS with dialogues. We also design an intuitive baseline model for comparison with existing state-of-the-art TTS methods on our leaderboard. Experimental results indicate the quality and effectiveness of DNASpeech, validating its potential to drive advancements in the TTS field. Dataset is available at https: //github.com/steven-ccq/DNASpeech .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing SynthesizersKai Shen, Zeqian Ju, Xu Tan, Eric Liu et al.ICLR 2024 · 362 citations
- Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech GenerationDongchan Min, Dong Bok Lee, Eunho Yang, Sung Ju HwangICML 2021 · 218 citations
- PromptTTS 2: Describing and Generating Voices with Text PromptYichong Leng, Zhifang Guo, Kai Shen, Zeqian Ju et al.ICLR 2024 · 80 citations
Related papers
- Emotionally Situated Text-to-Speech Synthesis in User-Agent ConversationYuchen Liu, Haoyu Zhang, Shichao Liu, Xiang Yin et al.ACM MM 2023 · 6 citations
- Scaling Rich Style-Prompted Text-to-Speech DatasetsAnuj Diwan, Zhisheng Zheng, David Harwath, Eunsol ChoiEMNLP 2025 · 2 citations
- Generative Expressive Conversational Speech SynthesisRui Liu, Yifan Hu, Yi Ren, Xiang Yin et al.ACM MM 2024 · 15 citations
- MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech SynthesisWenhao Guan, Yishuang Li, Tao Li, Hukai Huang et al.AAAI 2024 · 25 citations
- Paralinguistics-Aware Speech-Empowered Large Language Models for Natural ConversationHeeseung Kim, Soonshin Seo, Kyeongseok Jeong, Ohsung Kwon et al.NeurIPS 2024 · 28 citations
