OnomaCap: Making Non-speech Sound Captions Accessible and Enjoyable through Onomatopoeic Sound Representation
JooYeong Kim, Jin-Hyuk Hong
Abstract
Non-speech sounds play an important role in setting the mood of a video and aiding comprehension. However, current non-speech sound captioning practices focus primarily on sound categories, which fails to provide a rich sound experience for d/Deaf and hard-of-hearing (DHH) viewers. Onomatopoeia, which succinctly captures expressive sound information, offers a potential solution but remains underutilized in non-speech sound captioning. This paper investigates how onomatopoeia benefits DHH audiences in non-speech sound captioning. We collected 7,962 sound-onomatopoeia pairs from listeners and developed a sound-onomatopoeia model that automatically transcribes sounds into onomatopoeic descriptions indistinguishable from human-generated ones. A user evaluation of 25 DHH participants using the model-generated onomatopoeia demonstrated that onomatopoeia significantly improved their video viewing experience. Participants most favored captions with onomatopoeia and category, and expressed a desire to see such captions across genres. We discuss the benefits and challenges of using onomatopoeia in non-speech sound captions, offering insights for future practices.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bdad4145-4ed2-4398-b5d1-8d8e1fff2011Cited by top-tier papers1
Ask how each one uses itRelated papers
- Watch It, Don't Imagine It: Creating a Better Caption-Occlusion Metric by Collecting More Ecologically Valid Judgments from DHH ViewersAkhter Al Amin, Saad Hassan, Sooyeon Lee, Matt HuenerfauthCHI 2022 · 17 citations
- Haptic-Captioning: Using Audio-Haptic Interfaces to Enhance Speaker Indication in Real-Time Captions for Deaf and Hard-of-Hearing ViewersYiwen Wang, Ziming Li, Pratheep Kumar Chelladurai, Wendy Dannels et al.CHI 2023 · 32 citations
- Visible Nuances: A Caption System to Visualize Paralinguistic Speech Cues for Deaf and Hard-of-Hearing IndividualsJooYeong Kim, Sooyeon Ahn, Jin-Hyuk HongCHI 2023 · 24 citations
- Visualization of Speech Prosody and Emotion in Captions: Accessibility for Deaf and Hard-of-Hearing UsersCaluã de Lacerda Pataca, Matthew Watkins, Roshan L. Peiris, Sooyeon Lee et al.CHI 2023 · 39 citations
- Unspoken Sound: Identifying Trends in Non-Speech Audio Captioning on YouTubeLloyd May, Keita Ohshiro, Khang Dang, Sripathi Sridhar et al.CHI 2024 · 10 citations
