Understanding News Text and Images Connection with Context-enriched Multimodal Transformers
Cláudio Bartolomeu, Rui Nóbrega, David Semedo
Abstract
The connection between news and the images that illustrate them goes beyond visual concept to natural language matching. Instead, the open-domain and event-reporting nature of news leads to semantically complex texts, in which images are used as a contextualizing element. This connection is often governed by a certain level of indirection, with journalistic criteria also playing an important role. In this paper, we address the complex challenge of connecting images to news text. A context-enriched Multimodal Transformer model is proposed, NewsLXMERT, capable of jointly attending to complementary multimodal news data perspectives. The idea is to create knowledge-rich and diverse multimodal sequences, going beyond the news headline (often lacking the necessary context) and visual objects, to effectively ground images to news pieces.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Visual News: Benchmark and Challenges in News Image CaptioningFuxiao Liu, Yinghan Wang, Tianlu Wang, Vicente OrdonezEMNLP 2021 · 67 citations
- Transform and Tell: Entity-Aware News Image CaptioningAlasdair Tran, Alexander Patrick Mathews, Lexing XieCVPR 2020
- MuMuQA: Multimedia Multi-Hop News Question Answering via Cross-Media Knowledge Extraction and GroundingRevanth Gangi Reddy, Xilin Rui, Manling Li, Xudong Lin et al.AAAI 2022 · 37 citations
- Fine-tuning with Multi-modal Entity Prompts for News Image CaptioningJingjing Zhang, Shancheng Fang, Zhendong Mao, Zhiwei Zhang et al.ACM MM 2022 · 16 citations
- News Content Completion with Location-Aware Image SelectionZhengkun Zhang, Jun Wang, Adam Jatowt, Zhe Sun et al.AAAI 2021 · 2 citations
