Visualization of Speech Prosody and Emotion in Captions: Accessibility for Deaf and Hard-of-Hearing Users
Caluã de Lacerda Pataca, Matthew Watkins, Roshan L. Peiris, Sooyeon Lee, Matt Huenerfauth
Abstract
Speech is expressive in ways that caption text does not capture, with emotion or emphasis information not conveyed. We interviewed eight Deaf and Hard-of-Hearing (dhh) individuals to understand if and how captions’ inexpressiveness impacts them in online meetings with hearing peers. Automatically captioned speech, we found, lacks affective depth, lending it a hard-to-parse ambiguity and general dullness. Interviewees regularly feel excluded, which some understand is an inherent quality of these types of meetings rather than a consequence of current caption text design. Next, we developed three novel captioning models that depicted, beyond words, features from prosody, emotions, and a mix of both. In an empirical study, 16 dhh participants compared these models with conventional captions. The emotion-based model outperformed traditional captions in depicting emotions and emphasis, with only a moderate loss in legibility, suggesting its potential as a more inclusive design for captions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers8
- EmoWear: Exploring Emotional Teasers for Voice Message Interaction on SmartwatchesPengcheng An, Jiawen Stefanie Zhu, Zibo Zhang, Yifei Yin et al.CHI 2024 · 19 citations
- Unspoken Sound: Identifying Trends in Non-Speech Audio Captioning on YouTubeLloyd May, Keita Ohshiro, Khang Dang, Sripathi Sridhar et al.CHI 2024 · 10 citations
- DanModCap: Designing a Danmaku Moderation Tool for Video-Sharing Platforms that Leverages Impact Captions with Large Language ModelsSiying Hu, Huanchen Wang, Yu Zhang, Piaohong Wang et al.CSCW 2025 · 6 citations
- SpeechCap: Leveraging Playful Impact Captions to Facilitate Interpersonal Communication in Social Virtual RealityYu Zhang, Yi Wen, Siying Hu, Zhicong LuCSCW 2025 · 6 citations
- Fuzzy Feelings: Arousal's Interpretive Noise and the Case for Acoustic-Based HapticsCaluã de Lacerda Pataca, Stephanie Patterson, Roshan L. Peiris, Matt HuenerfauthCHI 2026 · 2 citations
Related papers
- Social, Environmental, and Technical: Factors at Play in the Current Use and Future Design of Small-Group CaptioningEmma J. McDonnell, Ping Liu, Steven M. Goodman, Raja S. Kushalnagar et al.CSCW 2021 · 50 citations
- Caption Royale: Exploring the Design Space of Affective Captions from the Perspective of Deaf and Hard-of-Hearing IndividualsCaluã de Lacerda Pataca, Saad Hassan, Nathan Tinker, Roshan Lalintha Peiris et al.CHI 2024 · 30 citations
- "Easier or Harder, Depending on Who the Hearing Person Is": Codesigning Videoconferencing Tools for Small Groups with Mixed Hearing StatusEmma J. McDonnell, Soo Hyun Moon, Lucy Jiang, Steven M. Goodman et al.CHI 2023 · 26 citations
- Understanding and Enhancing The Role of Speechreading in Online d/DHH Communication AccessibilityAashaka Desai, Jennifer Mankoff, Richard E. LadnerCHI 2023 · 9 citations
- Visible Nuances: A Caption System to Visualize Paralinguistic Speech Cues for Deaf and Hard-of-Hearing IndividualsJooYeong Kim, Sooyeon Ahn, Jin-Hyuk HongCHI 2023 · 24 citations
