Unspoken Sound: Identifying Trends in Non-Speech Audio Captioning on YouTube
Lloyd May, Keita Ohshiro, Khang Dang, Sripathi Sridhar, Jhanvi Pai, Magdalena Fuentes, Sooyeon Lee, Mark Cartwright
摘要
High-quality closed captioning of both speech and non-speech elements (e.g., music, sound effects, manner of speaking, and speaker identification) is essential for the accessibility of video content, especially for d/Deaf and hard-of-hearing individuals. While many regions have regulations mandating captioning for television and movies, a regulatory gap remains for the vast amount of web-based video content, including the staggering 500+ hours uploaded to YouTube every minute. Advances in automatic speech recognition have bolstered the presence of captions on YouTube. However, the technology has notable limitations, including the omission of many non-speech elements, which are often crucial for understanding content narratives. This paper examines the contemporary and historical state of non-speech information (NSI) captioning on YouTube through the creation and exploratory analysis of a dataset of over 715k videos. We identify factors that influence NSI caption practices and suggest avenues for future research to enhance the accessibility of online video content.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Fuzzy Feelings: Arousal's Interpretive Noise and the Case for Acoustic-Based HapticsCaluã de Lacerda Pataca, Stephanie Patterson, Roshan L. Peiris, Matt HuenerfauthCHI 2026 · 被引用 2 次
- Like, Comment & Caption: A Decade of Social Media Video Caption Research (2015-2025)Huong Nguyen, Emma J. McDonnell, Lloyd May, Alexander Druzenko 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper7
- Rescribe: Authoring and Automatically Editing Audio DescriptionsAmy Pavel, Gabriel Reyes, Jeffrey P. BighamUIST 2020 · 被引用 72 次
- AVscript: Accessible Video Editing with Audio-Visual ScriptsMina Huh, Saelyne Yang, Yi-Hao Peng, Xiang 'Anthony' Chen 等CHI 2023 · 被引用 44 次
- A View on the Viewer: Gaze-Adaptive Captions for VideosKuno Kurzhals, Fabian Göbel, Katrin Angerbauer, Michael Sedlmair 等CHI 2020 · 被引用 42 次
- Visualization of Speech Prosody and Emotion in Captions: Accessibility for Deaf and Hard-of-Hearing UsersCaluã de Lacerda Pataca, Matthew Watkins, Roshan L. Peiris, Sooyeon Lee 等CHI 2023 · 被引用 39 次
- An Exploration of Captioning Practices and Challenges of Individual Content Creators on YouTube for People with Hearing ImpairmentsFranklin Mingzhe Li, Cheng Lu, Zhicong Lu, Patrick Carrington 等CSCW 2022 · 被引用 34 次
相关 Paper
- OnomaCap: Making Non-speech Sound Captions Accessible and Enjoyable through Onomatopoeic Sound RepresentationJooYeong Kim, Jin-Hyuk HongCHI 2025 · 被引用 7 次
- "Caption It in an Accessible Way That Is Also Enjoyable": Characterizing User-Driven Captioning Practices on TikTokEmma J. McDonnell, Tessa Eagle, Pitch Sinlapanuntakul, Soo Hyun Moon 等CHI 2024 · 被引用 33 次
- Visible Nuances: A Caption System to Visualize Paralinguistic Speech Cues for Deaf and Hard-of-Hearing IndividualsJooYeong Kim, Sooyeon Ahn, Jin-Hyuk HongCHI 2023 · 被引用 24 次
- Spoken Moments: Learning Joint Audio-Visual Representations From Video DescriptionsMathew Monfort, SouYoung Jin, Alexander H. Liu, David Harwath 等CVPR 2021
- Toward a Multi-modal Understanding of Visual-Linguistic Design of Telop: A Computational Analysis of Telop Used in Korean YouTube VideosTaeyoung Ko, Kyungho LeeCSCW 2026
