Lune

ACL2026Top-tier venue

HowToNarrate: A General-Domain Benchmark for Synchronized Video Narration with External Knowledge

Xueyan Wang, Dingyi Yang, Qin Jin

2026Year

Abstract

We present HowToNarrate, the first generaldomain benchmark for Synchronized Video Narration. The benchmark contains 3.2K videos across seven domains, segmented into 37.5K clips with aligned narrations and associated external knowledge. Effective narration requires models to understand visual scenes, incorporate relevant knowledge, and produce coherent, length-appropriate descriptions. We systematically benchmark current Multimodal LLMs (MLLMs) on these abilities. Our analysis shows that existing MLLMs overemphasize knowledge retrieval while largely neglecting prior context (receiving less than 10% attention). Moreover, they often conflate narration context with external knowledge, leading to redundancy and incoherence. To mitigate these issues, we propose VideoNarrationAgent, a multi-agent framework that combines context compression, knowledge retrieval, and narration generation. Experiments demonstrate that our method significantly improves MLLM performance. Furthermore, instruction tuning on HowToNarrate enhances both context-awareness and length control, boosting Qwen2.5-VL's score from 25 to 84. Our dataset and codes are released at https://github. com/wangxueyan666/HowToNarrate .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 4b8bf50b-6b8c-43f2-9fa7-c70d7cb55bab

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines