What You See is What You Ask: Evaluating Audio Descriptions
Divy Kala, Eshika Khandelwal, Makarand Tapaswi
Abstract
Audio descriptions (ADs) narrate important visual details in movies, enabling Blind and Low Vision (BLV) users to understand narratives and appreciate visual details.Existing works in automatic AD generation mostly focus on few-second trimmed clips, and evaluate them by comparing against a single groundtruth reference AD.However, writing ADs is inherently subjective.Through alignment and analysis of two independent AD tracks for the same movies, we quantify the subjectivity in when and whether to describe, and what and how to highlight.Thus, we show that working with trimmed clips is inadequate.We propose ADQA, a QA benchmark that evaluates ADs at the level of few-minute long, coherent video segments, testing whether they would help BLV users understand the story and appreciate visual details.ADQA features visual appreciation (VA) questions about visual facts and narrative understanding (NU) questions based on the plot.Through ADQA, we show that current AD generation methods lag far behind humanauthored ADs.We conclude with several recommendations for future work and introduce a public leaderboard for benchmarking.Generate Question Generate Question Generate Answer . . . . . .Dialog: I wish that for only one day Dad couldn'
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fcb6f141-2e2d-4b20-97db-183b9d0ad1ccBuilds on9
- MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio DescriptionsMattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron et al.CVPR 2022 · 84 citations
- Rescribe: Authoring and Automatically Editing Audio DescriptionsAmy Pavel, Gabriel Reyes, Jeffrey P. BighamUIST 2020 · 72 citations
- MM-Narrator: Narrating Long-form Videos with Multimodal In-Context LearningChaoyi Zhang, Kevin Lin, Zhengyuan Yang, Jianfeng Wang et al.CVPR 2024 · 20 citations
- MICap: A Unified Model for Identity-Aware Movie DescriptionsHaran Raajesh, Naveen Reddy Desanur, Zeeshan Khan, Makarand TapaswiCVPR 2024 · 4 citations
- Contextual AD Narration with Interleaved Multimodal SequenceHanlin Wang, Zhan Tong, Kecheng Zheng, Yujun Shen et al.CVPR 2025
Related papers
- AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionTengda Han, Max Bain, Arsha Nagrani, Gül Varol et al.ICCV 2023 · 55 citations
- AutoAD III: The Prequel - Back to the PixelsTengda Han, Max Bain, Arsha Nagrani, Gül Varol et al.CVPR 2024
- Toward Automatic Audio Description Generation for Accessible VideosYujia Wang, Wei Liang, Haikun Huang, Yongqi Zhang et al.CHI 2021 · 86 citations
- DistinctAD: Distinctive Audio Description Generation in ContextsBo Fang, Wenhao Wu, Qiangqiang Wu, Yuxin Song et al.CVPR 2025
- VideoA11y: Method and Dataset for Accessible Video DescriptionChaoyu Li, Sid Padmanabhuni, Maryam S. Cheema, Hasti Seifi et al.CHI 2025 · 23 citations
