VIEWS: Entity-Aware News Video Captioning
Hammad A. Ayyubi, Tianqi Liu, Arsha Nagrani, Xudong Lin, Mingda Zhang, Anurag Arnab, Feng Han, Yukun Zhu, Xuande Feng, Kevin Zhang, Jialu Liu, Shih-Fu Chang
Abstract
Existing popular video captioning benchmarks and models often produce generic captions for videos that lack specific identification of individuals, locations, or organizations (namedentities). However, in the case of news videos, the setting is more demanding, requiring the inclusion of such named entities for meaningful summarization. Therefore, we introduce the task of directly summarizing news videos into captions that are entities-aware. To facilitate research in this area, we have collected a large-scale dataset named VIEWS (VIdeo NEWS). Within this task, we face challenges inherent to recognizing named entities and navigating diverse, dynamic contexts, all while relying solely on visual cues. To address these challenges, we propose a model-agnostic approach that enriches visual information extracted from videos with context sourced from external knowledge, enabling the generation of entity-aware captions. We validate the effectiveness of our approach across three video captioning models. Additionally, we conduct a critical analysis of our methodology to gain insights into the complexity of the task, the challenges it presents, and potential avenues for future research. WRAP US Pres arrives at next stop on tour of African nations ADDS presser 1. Wide of helicopter arriving carrying US President Bush 2. Mid of Bush getting out of helicopter with wife, Laura 3. Bush meeting Liberian officials… STORYLINE President George W. Bush said on Thursday that president of the … WRAP US Pres arrives at next stop on tour of African nations ADDS presser 1. Wide of helicopter arriving carrying US President Bush 2. Mid of Bush getting out of helicopter with wife, Laura … US President George W. Bush arrived in Liberia as part of his tour of African nations. He was greeted by Liberian President Ellen Johnson Sirleaf.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Cut to the Chase: Training-free Multimodal Summarization via Chain-of-EventsXiaoxing You, Qiang Huang, Lingyu Li, Xiaojun Chang et al.CVPR 2026 · 4 citations
- ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific ExperimentsJiali Chen, Yujie Jia, Zihan Wu, Jinyu Yang et al.ACM MM 2025 · 1 citation
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
Related papers
- Visual News: Benchmark and Challenges in News Image CaptioningFuxiao Liu, Yinghan Wang, Tianlu Wang, Vicente OrdonezEMNLP 2021 · 67 citations
- ICECAP: Information Concentrated Entity-aware Image CaptioningAnwen Hu, Shizhe Chen, Qin JinACM MM 2020 · 20 citations
- Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image CaptioningXiaoxing You, Qiang Huang, Lingyu Li, Chi Zhang et al.AAAI 2026 · 1 citation
- Discriminative Latent Semantic Graph for Video CaptioningYang Bai, Junyan Wang, Yang Long, Bingzhang Hu et al.ACM MM 2021 · 26 citations
- Journalistic Guidelines Aware News Image CaptioningXuewen Yang, Svebor Karaman, Joel R. Tetreault, Alejandro JaimesEMNLP 2021
