DPP-TTS: Diversifying prosodic features of speech via determinantal point processes
Seongho Joo, Hyukhun Koh, Kyomin Jung
Abstract
With the rapid advancement in deep generative models, recent neural Text-To-Speech (TTS) models have succeeded in synthesizing humanlike speech. There have been some efforts to generate speech with various prosody beyond monotonous prosody patterns. However, previous works have several limitations. First, typical TTS models depend on the scaled sampling temperature for boosting the diversity of prosody. Speech samples generated at high sampling temperatures often lack perceptual prosodic diversity, thereby hampering the naturalness of the speech. Second, the diversity among samples is neglected since the sampling procedure often focuses on a single speech sample rather than multiple ones. In this paper, we propose DPP-TTS: a text-to-speech model based on Determinantal Point Processes (DPPs) with a new objective function and prosody diversifying module. Our TTS model is capable of generating speech samples that simultaneously consider perceptual diversity in each sample and among multiple samples. We demonstrate that DPP-TTS generates speech samples with more diversified prosody than baselines in the side-by-side comparison test considering the naturalness of speech at the same time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1b28766-ca08-4b3f-b388-c7662d6460ffBuilds on11
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 663 citations
- DiffSinger: Singing Voice Synthesis via Shallow Diffusion MechanismJinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen et al.AAAI 2022 · 348 citations
Related papers
- Cross-Utterance Conditioned VAE for Non-Autoregressive Text-to-SpeechYang Li, Cheng Yu, Guangzhi Sun, Hua Jiang et al.ACL 2022 · 7 citations
- ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-SpeechRongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu et al.ACM MM 2022 · 182 citations
- Cauchy Diffusion: A Heavy-tailed Denoising Diffusion Probabilistic Model for Speech SynthesisQi Lian, Yu Qi, Yueming WangAAAI 2025 · 3 citations
- Diverse Video Generation with Determinantal Point Process-Guided Policy OptimizationTahira Kazimi, Connor Dunlop, Pinar YanardagCVPR 2026 · 4 citations
- Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech GenerationDongchan Min, Dong Bok Lee, Eunho Yang, Sung Ju HwangICML 2021 · 218 citations
