SpeakEasy: Enhancing Text-to-Speech Interactions for Expressive Content Creation
Stephen Brade, Sam Anderson, Rithesh Kumar, Zeyu Jin, Anh Truong
Abstract
Novice content creators often invest significant time recording expressive speech for social media videos. While recent advancements in text-to-speech (TTS) technology can generate highly realistic speech in various languages and accents, many struggle with unintuitive or overly granular TTS interfaces. We propose simplifying TTS generation by allowing users to specify high-level context alongside their script. Our Wizard-of-Oz system, SpeakEasy, leverages user-provided context to inform and influence TTS output, enabling iterative refinement with high-level feedback. This approach was informed by two 8-subject formative studies: one examining content creators' experiences with TTS, and the other drawing on effective strategies from voice actors. Our evaluation shows that participants using SpeakEasy were more successful in generating performances matching their personal standards, without requiring significantly more effort than leading industry interfaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d01fd088-d46f-491f-a177-147a560b2719Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language ModelsStephen Brade, Bryan Wang, Maurício Sousa, Sageev Oore et al.UIST 2023 · 179 citations
- PromptTTS 2: Describing and Generating Voices with Text PromptYichong Leng, Zhifang Guo, Kai Shen, Zeqian Ju et al.ICLR 2024 · 80 citations
- Choice of Voices: A Large-Scale Evaluation of Text-to-Speech Voice Quality for Long-Form ContentJulia Cambre, Jessica Colnago, Jim Maddock, Janice Y. Tsai et al.CHI 2020 · 65 citations
- Smartphone Usage by Expert Blind UsersMohit Jain, Nirmalendu Diwakar, Manohar SwaminathanCHI 2021 · 40 citations
- VoiceCoach: Interactive Evidence-based Training for Voice Modulation Skills in Public SpeakingXingbo Wang, Haipeng Zeng, Yong Wang, Aoyu Wu et al.CHI 2020 · 34 citations
Related papers
- AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-SpeechBin Kang, Shaoguo Wen, Yang Fan, Shunlong Wu et al.ICML 2026
- Speech Recognition Model Improves Text-to-Speech Synthesis Using Fine-Grained RewardGuansu Wang, Peijie SunAAAI 2026
- Eliciting Implicit Acoustic Styles from Open-domain Instructions to Facilitate Fine-grained Controllable Generation of SpeechJianxing Yu, Zihao Gou, Chen Li, Zhisheng Wang et al.EMNLP 2025
- SpeechCraft: A Fine-Grained Expressive Speech Dataset with Natural Language DescriptionZeyu Jin, Jia Jia, Qixin Wang, Kehan Li et al.ACM MM 2024 · 12 citations
- Emotionally Situated Text-to-Speech Synthesis in User-Agent ConversationYuchen Liu, Haoyu Zhang, Shichao Liu, Xiang Yin et al.ACM MM 2023 · 6 citations
