Toward Automatic Audio Description Generation for Accessible Videos
Yujia Wang, Wei Liang, Haikun Huang, Yongqi Zhang, Dingzeyu Li, Lap-Fai Yu
摘要
Video accessibility is essential for people with visual impairments. Audio descriptions describe what is happening on-screen, e.g., physical actions, facial expressions, and scene changes. Generating high-quality audio descriptions requires a lot of manual description generation [50]. To address this accessibility obstacle, we built a system that analyzes the audiovisual contents of a video and generates the audio descriptions. The system consisted of three modules: AD insertion time prediction, AD generation, and AD optimization. We evaluated the quality of our system on five types of videos by conducting qualitative studies with 20 sighted users and 12 users who were blind or visually impaired. Our findings revealed how audio description preferences varied with user types and video types. Based on our study’s analysis, we provided recommendations for the development of future audio description generation technologies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- A Literature Review of Video-Sharing Platform Research in HCIAva Bartolome, Shuo NiuCHI 2023 · 被引用 63 次
- AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionTengda Han, Max Bain, Arsha Nagrani, Gül Varol 等ICCV 2023 · 被引用 55 次
- Making Short-Form Videos Accessible with Hierarchical Video SummariesTess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C. Derry 等CHI 2024 · 被引用 37 次
- SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision ViewersZheng Ning, Brianna L. Wimer, Kaiwen Jiang, Keyi Chen 等CHI 2024 · 被引用 27 次
- DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio DiscussionsShuchang Xu, Xiaofu Jin, Huamin Qu, Yukang YanCHI 2025 · 被引用 26 次
它引用的顶会 Paper3
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 被引用 992 次
- Twitter A11y: A Browser Extension to Make Twitter Images AccessibleCole Gleason, Amy Pavel, Emma McCamey, Christina Low 等CHI 2020 · 被引用 123 次
- Scene-Aware Background Music SynthesisYujia Wang, Wei Liang, Wanwan Li, Dingzeyu Li 等ACM MM 2020 · 被引用 14 次
相关 Paper
- What Makes Videos Accessible to Blind and Visually Impaired People?Xingyu Liu, Patrick Carrington, Xiang 'Anthony' Chen, Amy PavelCHI 2021 · 被引用 78 次
- Rescribe: Authoring and Automatically Editing Audio DescriptionsAmy Pavel, Gabriel Reyes, Jeffrey P. BighamUIST 2020 · 被引用 72 次
- What You See is What You Ask: Evaluating Audio DescriptionsDivy Kala, Eshika Khandelwal, Makarand TapaswiEMNLP 2025
- Supporting Novices Author Audio Descriptions via Automatic FeedbackRosiana Natalie, Joshua Tseng, Hernisa Kacorri, Kotaro HaraCHI 2023 · 被引用 18 次
- CrossA11y: Identifying Video Accessibility Issues via Cross-modal GroundingXingyu Bruce Liu, Ruolin Wang, Dingzeyu Li, Xiang Anthony Chen 等UIST 2022 · 被引用 24 次
