Understanding News Text and Images Connection with Context-enriched Multimodal Transformers
Cláudio Bartolomeu, Rui Nóbrega, David Semedo
摘要
The connection between news and the images that illustrate them goes beyond visual concept to natural language matching. Instead, the open-domain and event-reporting nature of news leads to semantically complex texts, in which images are used as a contextualizing element. This connection is often governed by a certain level of indirection, with journalistic criteria also playing an important role. In this paper, we address the complex challenge of connecting images to news text. A context-enriched Multimodal Transformer model is proposed, NewsLXMERT, capable of jointly attending to complementary multimodal news data perspectives. The idea is to create knowledge-rich and diverse multimodal sequences, going beyond the news headline (often lacking the necessary context) and visual objects, to effectively ground images to news pieces.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Visual News: Benchmark and Challenges in News Image CaptioningFuxiao Liu, Yinghan Wang, Tianlu Wang, Vicente OrdonezEMNLP 2021 · 被引用 67 次
- Transform and Tell: Entity-Aware News Image CaptioningAlasdair Tran, Alexander Patrick Mathews, Lexing XieCVPR 2020
- MuMuQA: Multimedia Multi-Hop News Question Answering via Cross-Media Knowledge Extraction and GroundingRevanth Gangi Reddy, Xilin Rui, Manling Li, Xudong Lin 等AAAI 2022 · 被引用 37 次
- Fine-tuning with Multi-modal Entity Prompts for News Image CaptioningJingjing Zhang, Shancheng Fang, Zhendong Mao, Zhiwei Zhang 等ACM MM 2022 · 被引用 16 次
- News Content Completion with Location-Aware Image SelectionZhengkun Zhang, Jun Wang, Adam Jatowt, Zhe Sun 等AAAI 2021 · 被引用 2 次
