Visualization of Speech Prosody and Emotion in Captions: Accessibility for Deaf and Hard-of-Hearing Users
Caluã de Lacerda Pataca, Matthew Watkins, Roshan L. Peiris, Sooyeon Lee, Matt Huenerfauth
摘要
Speech is expressive in ways that caption text does not capture, with emotion or emphasis information not conveyed. We interviewed eight Deaf and Hard-of-Hearing (dhh) individuals to understand if and how captions’ inexpressiveness impacts them in online meetings with hearing peers. Automatically captioned speech, we found, lacks affective depth, lending it a hard-to-parse ambiguity and general dullness. Interviewees regularly feel excluded, which some understand is an inherent quality of these types of meetings rather than a consequence of current caption text design. Next, we developed three novel captioning models that depicted, beyond words, features from prosody, emotions, and a mix of both. In an empirical study, 16 dhh participants compared these models with conventional captions. The emotion-based model outperformed traditional captions in depicting emotions and emphasis, with only a moderate loss in legibility, suggesting its potential as a more inclusive design for captions.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper8
- EmoWear: Exploring Emotional Teasers for Voice Message Interaction on SmartwatchesPengcheng An, Jiawen Stefanie Zhu, Zibo Zhang, Yifei Yin 等CHI 2024 · 被引用 19 次
- Unspoken Sound: Identifying Trends in Non-Speech Audio Captioning on YouTubeLloyd May, Keita Ohshiro, Khang Dang, Sripathi Sridhar 等CHI 2024 · 被引用 10 次
- DanModCap: Designing a Danmaku Moderation Tool for Video-Sharing Platforms that Leverages Impact Captions with Large Language ModelsSiying Hu, Huanchen Wang, Yu Zhang, Piaohong Wang 等CSCW 2025 · 被引用 6 次
- SpeechCap: Leveraging Playful Impact Captions to Facilitate Interpersonal Communication in Social Virtual RealityYu Zhang, Yi Wen, Siying Hu, Zhicong LuCSCW 2025 · 被引用 6 次
- Fuzzy Feelings: Arousal's Interpretive Noise and the Case for Acoustic-Based HapticsCaluã de Lacerda Pataca, Stephanie Patterson, Roshan L. Peiris, Matt HuenerfauthCHI 2026 · 被引用 2 次
相关 Paper
- Social, Environmental, and Technical: Factors at Play in the Current Use and Future Design of Small-Group CaptioningEmma J. McDonnell, Ping Liu, Steven M. Goodman, Raja S. Kushalnagar 等CSCW 2021 · 被引用 50 次
- Caption Royale: Exploring the Design Space of Affective Captions from the Perspective of Deaf and Hard-of-Hearing IndividualsCaluã de Lacerda Pataca, Saad Hassan, Nathan Tinker, Roshan Lalintha Peiris 等CHI 2024 · 被引用 30 次
- "Easier or Harder, Depending on Who the Hearing Person Is": Codesigning Videoconferencing Tools for Small Groups with Mixed Hearing StatusEmma J. McDonnell, Soo Hyun Moon, Lucy Jiang, Steven M. Goodman 等CHI 2023 · 被引用 26 次
- Understanding and Enhancing The Role of Speechreading in Online d/DHH Communication AccessibilityAashaka Desai, Jennifer Mankoff, Richard E. LadnerCHI 2023 · 被引用 9 次
- Visible Nuances: A Caption System to Visualize Paralinguistic Speech Cues for Deaf and Hard-of-Hearing IndividualsJooYeong Kim, Sooyeon Ahn, Jin-Hyuk HongCHI 2023 · 被引用 24 次
