Lune

EMNLP2025Top-tier venue

What You See is What You Ask: Evaluating Audio Descriptions

Divy Kala, Eshika Khandelwal, Makarand Tapaswi

2025Year

Abstract

Audio descriptions (ADs) narrate important visual details in movies, enabling Blind and Low Vision (BLV) users to understand narratives and appreciate visual details.Existing works in automatic AD generation mostly focus on few-second trimmed clips, and evaluate them by comparing against a single groundtruth reference AD.However, writing ADs is inherently subjective.Through alignment and analysis of two independent AD tracks for the same movies, we quantify the subjectivity in when and whether to describe, and what and how to highlight.Thus, we show that working with trimmed clips is inadequate.We propose ADQA, a QA benchmark that evaluates ADs at the level of few-minute long, coherent video segments, testing whether they would help BLV users understand the story and appreciate visual details.ADQA features visual appreciation (VA) questions about visual facts and narrative understanding (NU) questions based on the plot.Through ADQA, we show that current AD generation methods lag far behind humanauthored ADs.We conclude with several recommendations for future work and introduce a public leaderboard for benchmarking.Generate Question Generate Question Generate Answer . . . . . .Dialog: I wish that for only one day Dad couldn'

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext fcb6f141-2e2d-4b20-97db-183b9d0ad1cc

Builds on9

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines