Choice of Voices: A Large-Scale Evaluation of Text-to-Speech Voice Quality for Long-Form Content
Julia Cambre, Jessica Colnago, Jim Maddock, Janice Y. Tsai, Jofish Kaye
摘要
The advancement of text-to-speech (TTS) voices and a rise of commercial TTS platforms allow people to easily experience TTS voices across a variety of technologies, applications, and form factors. As such, we evaluated TTS voices for long-form content: not individual words or sentences, but voices that are pleasant to listen to for several minutes at a time. We introduce a method using a crowdsourcing platform and an online survey to evaluate voices based on listening experience, perception of clarity and quality, and comprehension. We evaluated 18 TTS voices, three human voices, and a text-only control condition. We found that TTS voices are close to rivaling human voices, yet no single voice outperforms the others across all evaluation dimensions. We conclude with considerations for selecting text-to-speech voices for long-form content.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper12
- Exploring Interactions Between Trust, Anthropomorphism, and Relationship Development in Voice AssistantsWilliam Seymour, Max Van KleekCSCW 2021 · 被引用 91 次
- Cruising Queer HCI on the DL: A Literature Review of LGBTQ+ People in HCIJordan Taylor, Ellen Simpson, Anh-Ton Tran, Jed R. Brubaker 等CHI 2024 · 被引用 55 次
- Social Media through Voice: Synthesized Voice Qualities and Self-presentationLotus Zhang, Lucy Jiang, Nicole Washington, Augustina Ao Liu 等CSCW 2021 · 被引用 40 次
- Firefox Voice: An Open and Extensible Voice Assistant Built Upon the WebJulia Cambre, Alex C. Williams, Afsaneh Razi, Ian Bicking 等CHI 2021 · 被引用 38 次
- KinVoices: Using Voices of Friends and Family in Voice InterfacesSamantha W. T. Chan, Tamil Selvan Gunasekaran, Yun Suen Pai, Haimo Zhang 等CSCW 2021 · 被引用 32 次
相关 Paper
- Long-Form Speech Generation with Spoken Language ModelsSe Jin Park, Julian Salazar, Aren Jansen, Keisuke Kinoshita 等ICML 2025
- SLVMEval: Synthetic Meta Evaluation Benchmark for Text-to-Long Video GenerationRyosuke Matsuda, Keito Kudo, Haruto Yoshida, Nobuyuki Shimizu 等CVPR 2026 · 被引用 2 次
- Hear Me Out: A Study on the Use of the Voice Modality for Crowdsourced Relevance AssessmentsNirmal Roy, Agathe Balayn, David Maxwell, Claudia HauffSIGIR 2023 · 被引用 1 次
- TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech SystemsChristoph Minixhofer, Ondrej Klejch, Peter BellICLR 2026 · 被引用 16 次
- "Hi! I am the Crowd Tasker" Crowdsourcing through Digital Voice AssistantsDanula Hettiachchi, Zhanna Sarsenbayeva, Fraser Allison, Niels van Berkel 等CHI 2020 · 被引用 22 次
