MM-Narrator: Narrating Long-form Videos with Multimodal In-Context Learning
Chaoyi Zhang, Kevin Lin, Zhengyuan Yang, Jianfeng Wang, Linjie Li, Chung-Ching Lin, Zicheng Liu, Lijuan Wang
Abstract
We present MM-Narrator, a novel system leveraging GPT-4 with multimodal in-context learning for the gener-ation of audio descriptions (AD). Unlike previous methods that primarily focused on downstream fine-tuning with short video clips, MM-Narrator excels in generating precise audio descriptions for videos of extensive lengths, even be-yond hours, in an autoregressive manner. This capability is made possible by the proposed memory-augmented generation process, which effectively utilizes both the short-term textual context and long-term visual memory through an efficient register-and-recall mechanism. These contextual memories compile pertinent past information, including storylines and character identities, ensuring an accurate tracking and depicting of story-coherent and character-centric audio descriptions. Maintaining the training-free design of MM-Narrator, we further propose a complexity-based demonstration selection strategy to largely enhance its multi-step reasoning capability via few-shot multimodal in-context learning (MM-ICL). Experimental results on MAD-eval dataset demonstrate that MM-Narrator consistently outperforms both the existing fine-tuning-based approaches and LLM-based approaches in most scenarios, as measured by standard evaluation metrics. Additionally, we introduce the first segment-based evaluator for recurrent text generation. Empowered by GPT-4, this evaluator comprehensively reasons and marks AD generation performance in various extendable dimensions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term MemoryLin Long, Yichen He, Wentao Ye, Yiyuan Pan et al.ICLR 2026 · 90 citations
- "It's Kind of Context Dependent": Understanding Blind and Low Vision People's Video Accessibility Preferences Across Viewing ScenariosLucy Jiang, Crescentia Jung, Mahika Phutane, Abigale Stangl et al.CHI 2024 · 23 citations
- Bringing RNNs Back to Efficient Open-Ended Video UnderstandingWeili Xu, Enxin Song, Wenhao Chai, Xuexiang Wen et al.ICCV 2025 · 12 citations
- Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision EncodersAli Rasekh, Erfan Bagheri Soula, Omid Daliran, Simon Gottschalk et al.NeurIPS 2025 · 10 citations
- TraveLER: A Modular Multi-LMM Agent Framework for Video Question-AnsweringChuyi Shang, Amos You, Sanjay Subramanian, Trevor Darrell et al.EMNLP 2024 · 7 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
Related papers
- DistinctAD: Distinctive Audio Description Generation in ContextsBo Fang, Wenhao Wu, Qiangqiang Wu, Yuxin Song et al.CVPR 2025
- HowToNarrate: A General-Domain Benchmark for Synchronized Video Narration with External KnowledgeXueyan Wang, Dingyi Yang, Qin JinACL 2026
- WorldMM: Dynamic Multimodal Memory Agent for Long Video ReasoningWoongyeong Yeo, Kangsan Kim, Jaehong Yoon, Sung Ju HwangCVPR 2026 · 53 citations
- JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and GenerationKai Liu, Jungang Li, Yuchong Sun, Shengqiong Wu et al.NeurIPS 2025 · 18 citations
- AutoAD III: The Prequel - Back to the PixelsTengda Han, Max Bain, Arsha Nagrani, Gül Varol et al.CVPR 2024
