Choice of Voices: A Large-Scale Evaluation of Text-to-Speech Voice Quality for Long-Form Content
Julia Cambre, Jessica Colnago, Jim Maddock, Janice Y. Tsai, Jofish Kaye
Abstract
The advancement of text-to-speech (TTS) voices and a rise of commercial TTS platforms allow people to easily experience TTS voices across a variety of technologies, applications, and form factors. As such, we evaluated TTS voices for long-form content: not individual words or sentences, but voices that are pleasant to listen to for several minutes at a time. We introduce a method using a crowdsourcing platform and an online survey to evaluate voices based on listening experience, perception of clarity and quality, and comprehension. We evaluated 18 TTS voices, three human voices, and a text-only control condition. We found that TTS voices are close to rivaling human voices, yet no single voice outperforms the others across all evaluation dimensions. We conclude with considerations for selecting text-to-speech voices for long-form content.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9c080859-e099-46a2-9815-55fd4759a1cbCited by top-tier papers12
- Exploring Interactions Between Trust, Anthropomorphism, and Relationship Development in Voice AssistantsWilliam Seymour, Max Van KleekCSCW 2021 · 91 citations
- Cruising Queer HCI on the DL: A Literature Review of LGBTQ+ People in HCIJordan Taylor, Ellen Simpson, Anh-Ton Tran, Jed R. Brubaker et al.CHI 2024 · 55 citations
- Social Media through Voice: Synthesized Voice Qualities and Self-presentationLotus Zhang, Lucy Jiang, Nicole Washington, Augustina Ao Liu et al.CSCW 2021 · 40 citations
- Firefox Voice: An Open and Extensible Voice Assistant Built Upon the WebJulia Cambre, Alex C. Williams, Afsaneh Razi, Ian Bicking et al.CHI 2021 · 38 citations
- KinVoices: Using Voices of Friends and Family in Voice InterfacesSamantha W. T. Chan, Tamil Selvan Gunasekaran, Yun Suen Pai, Haimo Zhang et al.CSCW 2021 · 32 citations
Related papers
- Long-Form Speech Generation with Spoken Language ModelsSe Jin Park, Julian Salazar, Aren Jansen, Keisuke Kinoshita et al.ICML 2025
- SLVMEval: Synthetic Meta Evaluation Benchmark for Text-to-Long Video GenerationRyosuke Matsuda, Keito Kudo, Haruto Yoshida, Nobuyuki Shimizu et al.CVPR 2026 · 2 citations
- Hear Me Out: A Study on the Use of the Voice Modality for Crowdsourced Relevance AssessmentsNirmal Roy, Agathe Balayn, David Maxwell, Claudia HauffSIGIR 2023 · 1 citation
- TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech SystemsChristoph Minixhofer, Ondrej Klejch, Peter BellICLR 2026 · 16 citations
- "Hi! I am the Crowd Tasker" Crowdsourcing through Digital Voice AssistantsDanula Hettiachchi, Zhanna Sarsenbayeva, Fraser Allison, Niels van Berkel et al.CHI 2020 · 22 citations
