What You See is What You Ask: Evaluating Audio Descriptions
Divy Kala, Eshika Khandelwal, Makarand Tapaswi
摘要
Audio descriptions (ADs) narrate important visual details in movies, enabling Blind and Low Vision (BLV) users to understand narratives and appreciate visual details.Existing works in automatic AD generation mostly focus on few-second trimmed clips, and evaluate them by comparing against a single groundtruth reference AD.However, writing ADs is inherently subjective.Through alignment and analysis of two independent AD tracks for the same movies, we quantify the subjectivity in when and whether to describe, and what and how to highlight.Thus, we show that working with trimmed clips is inadequate.We propose ADQA, a QA benchmark that evaluates ADs at the level of few-minute long, coherent video segments, testing whether they would help BLV users understand the story and appreciate visual details.ADQA features visual appreciation (VA) questions about visual facts and narrative understanding (NU) questions based on the plot.Through ADQA, we show that current AD generation methods lag far behind humanauthored ADs.We conclude with several recommendations for future work and introduce a public leaderboard for benchmarking.Generate Question Generate Question Generate Answer . . . . . .Dialog: I wish that for only one day Dad couldn'
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio DescriptionsMattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron 等CVPR 2022 · 被引用 84 次
- Rescribe: Authoring and Automatically Editing Audio DescriptionsAmy Pavel, Gabriel Reyes, Jeffrey P. BighamUIST 2020 · 被引用 72 次
- MM-Narrator: Narrating Long-form Videos with Multimodal In-Context LearningChaoyi Zhang, Kevin Lin, Zhengyuan Yang, Jianfeng Wang 等CVPR 2024 · 被引用 20 次
- MICap: A Unified Model for Identity-Aware Movie DescriptionsHaran Raajesh, Naveen Reddy Desanur, Zeeshan Khan, Makarand TapaswiCVPR 2024 · 被引用 4 次
- Contextual AD Narration with Interleaved Multimodal SequenceHanlin Wang, Zhan Tong, Kecheng Zheng, Yujun Shen 等CVPR 2025
相关 Paper
- AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionTengda Han, Max Bain, Arsha Nagrani, Gül Varol 等ICCV 2023 · 被引用 55 次
- AutoAD III: The Prequel - Back to the PixelsTengda Han, Max Bain, Arsha Nagrani, Gül Varol 等CVPR 2024
- Toward Automatic Audio Description Generation for Accessible VideosYujia Wang, Wei Liang, Haikun Huang, Yongqi Zhang 等CHI 2021 · 被引用 86 次
- DistinctAD: Distinctive Audio Description Generation in ContextsBo Fang, Wenhao Wu, Qiangqiang Wu, Yuxin Song 等CVPR 2025
- VideoA11y: Method and Dataset for Accessible Video DescriptionChaoyu Li, Sid Padmanabhuni, Maryam S. Cheema, Hasti Seifi 等CHI 2025 · 被引用 23 次
