Contextual AD Narration with Interleaved Multimodal Sequence
Hanlin Wang, Zhan Tong, Kecheng Zheng, Yujun Shen, Limin Wang
Abstract
The Audio Description (AD) task aims to generate descriptions of visual elements for visually impaired individuals to help them access long-form video content, like movies. With video feature, text, character bank and context information as inputs, the generated ADs are able to correspond to the characters by name and provide reasonable, contextual descriptions to help audience understand the storyline of movie. To achieve this goal, we propose to leverage pre-trained foundation models through a simple and unified framework to generate ADs with interleaved multimodal sequence as input, termed as Uni-AD. To enhance the alignment of features across various modalities with finer granularity, we introduce a simple and lightweight module that maps video features into the textual feature space. Moreover, we also propose a character-refinement module to provide more precise information by identifying the main characters who play more significant roles in the video context. With these unique designs, we further incorporate contextual information and a contrastive loss into our architecture to generate smoother and more contextually appropriate ADs. Experiments on multiple AD datasets show that Uni-AD performs well on AD generation, which demonstrates the effectiveness of our approach. Our code is available at: https://github.com/ant-research/UniAD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00c67f67-36f6-4fe7-8d42-6828f50d3a1eCited by top-tier papers5
- UniM: A Unified Any-to-Any Interleaved Multimodal BenchmarkYanlin Li, Minghui Guo, Kaiwen Zhang, Shize Zhang et al.CVPR 2026 · 10 citations
- DistinctAD: Distinctive Audio Description Generation in ContextsBo Fang, Wenhao Wu, Qiangqiang Wu, Yuxin Song et al.CVPR 2025
- Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description GenerationJunyu Xie, Tengda Han, Max Bain, Arsha Nagrani et al.ICCV 2025
- Character-Centric Understanding of Animated MoviesZhongrui Gui, Junyu Xie, Tengda Han, Weidi Xie et al.ACM MM 2025
- What You See is What You Ask: Evaluating Audio DescriptionsDivy Kala, Eshika Khandelwal, Makarand TapaswiEMNLP 2025
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- DreamLLM: Synergistic Multimodal Comprehension and CreationRunpei Dong, Chunrui Han, Yuang Peng, Zekun Qi et al.ICLR 2024 · 315 citations
Related papers
- AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionTengda Han, Max Bain, Arsha Nagrani, Gül Varol et al.ICCV 2023 · 55 citations
- AutoAD III: The Prequel - Back to the PixelsTengda Han, Max Bain, Arsha Nagrani, Gül Varol et al.CVPR 2024
- AutoAD: Movie Description in ContextTengda Han, Max Bain, Arsha Nagrani, Gül Varol et al.CVPR 2023
- Tell What You Hear From What You See - Video to Audio Generation Through TextXiulong Liu, Kun Su, Eli ShlizermanNeurIPS 2024 · 46 citations
- Auto-ACD: A Large-scale Dataset for Audio-Language Representation LearningLuoyi Sun, Xuenan Xu, Mengyue Wu, Weidi XieACM MM 2024 · 23 citations
