Unspoken Sound: Identifying Trends in Non-Speech Audio Captioning on YouTube
Lloyd May, Keita Ohshiro, Khang Dang, Sripathi Sridhar, Jhanvi Pai, Magdalena Fuentes, Sooyeon Lee, Mark Cartwright
Abstract
High-quality closed captioning of both speech and non-speech elements (e.g., music, sound effects, manner of speaking, and speaker identification) is essential for the accessibility of video content, especially for d/Deaf and hard-of-hearing individuals. While many regions have regulations mandating captioning for television and movies, a regulatory gap remains for the vast amount of web-based video content, including the staggering 500+ hours uploaded to YouTube every minute. Advances in automatic speech recognition have bolstered the presence of captions on YouTube. However, the technology has notable limitations, including the omission of many non-speech elements, which are often crucial for understanding content narratives. This paper examines the contemporary and historical state of non-speech information (NSI) captioning on YouTube through the creation and exploratory analysis of a dataset of over 715k videos. We identify factors that influence NSI caption practices and suggest avenues for future research to enhance the accessibility of online video content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8dfd28e1-a36b-4cac-8bb3-ed797981e013Cited by top-tier papers2
- Fuzzy Feelings: Arousal's Interpretive Noise and the Case for Acoustic-Based HapticsCaluã de Lacerda Pataca, Stephanie Patterson, Roshan L. Peiris, Matt HuenerfauthCHI 2026 · 2 citations
- Like, Comment & Caption: A Decade of Social Media Video Caption Research (2015-2025)Huong Nguyen, Emma J. McDonnell, Lloyd May, Alexander Druzenko et al.CHI 2026 · 1 citation
Builds on7
- Rescribe: Authoring and Automatically Editing Audio DescriptionsAmy Pavel, Gabriel Reyes, Jeffrey P. BighamUIST 2020 · 72 citations
- AVscript: Accessible Video Editing with Audio-Visual ScriptsMina Huh, Saelyne Yang, Yi-Hao Peng, Xiang 'Anthony' Chen et al.CHI 2023 · 44 citations
- A View on the Viewer: Gaze-Adaptive Captions for VideosKuno Kurzhals, Fabian Göbel, Katrin Angerbauer, Michael Sedlmair et al.CHI 2020 · 42 citations
- Visualization of Speech Prosody and Emotion in Captions: Accessibility for Deaf and Hard-of-Hearing UsersCaluã de Lacerda Pataca, Matthew Watkins, Roshan L. Peiris, Sooyeon Lee et al.CHI 2023 · 39 citations
- An Exploration of Captioning Practices and Challenges of Individual Content Creators on YouTube for People with Hearing ImpairmentsFranklin Mingzhe Li, Cheng Lu, Zhicong Lu, Patrick Carrington et al.CSCW 2022 · 34 citations
Related papers
- OnomaCap: Making Non-speech Sound Captions Accessible and Enjoyable through Onomatopoeic Sound RepresentationJooYeong Kim, Jin-Hyuk HongCHI 2025 · 7 citations
- "Caption It in an Accessible Way That Is Also Enjoyable": Characterizing User-Driven Captioning Practices on TikTokEmma J. McDonnell, Tessa Eagle, Pitch Sinlapanuntakul, Soo Hyun Moon et al.CHI 2024 · 33 citations
- Visible Nuances: A Caption System to Visualize Paralinguistic Speech Cues for Deaf and Hard-of-Hearing IndividualsJooYeong Kim, Sooyeon Ahn, Jin-Hyuk HongCHI 2023 · 24 citations
- Spoken Moments: Learning Joint Audio-Visual Representations From Video DescriptionsMathew Monfort, SouYoung Jin, Alexander H. Liu, David Harwath et al.CVPR 2021
- Toward a Multi-modal Understanding of Visual-Linguistic Design of Telop: A Computational Analysis of Telop Used in Korean YouTube VideosTaeyoung Ko, Kyungho LeeCSCW 2026
