Transform and Tell: Entity-Aware News Image Captioning
Alasdair Tran, Alexander Patrick Mathews, Lexing Xie
Abstract
We propose an end-to-end model which generates captions for images embedded in news articles. News images present two key challenges: they rely on real-world knowledge, especially about named entities; and they typically have linguistically rich captions that include uncommon words. We address the first challenge by associating words in the caption with faces and objects in the image, via a multi-modal, multi-head attention mechanism. We tackle the second challenge with a state-of-the-art transformer language model that uses byte-pair-encoding to generate captions as a sequence of word parts. On the Good-News dataset [3], our model outperforms the previous state of the art by a factor of four in CIDEr score (13 → 54). This performance gain comes from a unique combination of language models, word representation, image embeddings, face embeddings, object embeddings, and improvements in neural network design. We also introduce the NYTimes800k dataset which is 70% larger than GoodNews, has higher article quality, and includes the locations of images within articles as an additional contextual cue.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- Exploiting BERT for Multimodal Target Sentiment Classification through Input Space TranslationZaid Khan, Yun FuACM MM 2021 · 192 citations
- WikiDiverse: A Multimodal Entity Linking Dataset with Diversified Contextual Topics and Entity TypesXuwu Wang, Junfeng Tian, Min Gui, Zhixu Li et al.ACL 2022 · 75 citations
- Visual News: Benchmark and Challenges in News Image CaptioningFuxiao Liu, Yinghan Wang, Tianlu Wang, Vicente OrdonezEMNLP 2021 · 67 citations
- EI-CLIP: Entity-aware Interventional Contrastive Learning for E-commerce Cross-modal RetrievalHaoyu Ma, Handong Zhao, Zhe Lin, Ajinkya Kale et al.CVPR 2022 · 56 citations
- Automatic Radiology Reports Generation via Memory Alignment NetworkHongyu Shen, Mingtao Pei, Juncai Liu, Zhaoxing TianAAAI 2024 · 40 citations
Builds on2
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- On Extractive and Abstractive Neural Document Summarization with Transformer Language ModelsJonathan Pilault, Raymond Li, Sandeep Subramanian, Chris PalEMNLP 2020 · 186 citations
Related papers
- Fine-tuning with Multi-modal Entity Prompts for News Image CaptioningJingjing Zhang, Shancheng Fang, Zhendong Mao, Zhiwei Zhang et al.ACM MM 2022 · 16 citations
- Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image CaptioningXiaoxing You, Qiang Huang, Lingyu Li, Chi Zhang et al.AAAI 2026 · 1 citation
- ICECAP: Information Concentrated Entity-aware Image CaptioningAnwen Hu, Shizhe Chen, Qin JinACM MM 2020 · 20 citations
- End-to-End Transformer Based Model for Image CaptioningYiyu Wang, Jungang Xu, Yingfei SunAAAI 2022 · 178 citations
- Journalistic Guidelines Aware News Image CaptioningXuewen Yang, Svebor Karaman, Joel R. Tetreault, Alejandro JaimesEMNLP 2021
