MERL: Multimodal Event Representation Learning in Heterogeneous Embedding Spaces
Linhai Zhang, Deyu Zhou, Yulan He, Zeng Yang
摘要
Previous work has shown the effectiveness of using event representations for tasks such as script event prediction and stock market prediction. It is however still challenging to learn the subtle semantic differences between events based solely on textual descriptions of events often represented as (subject, predicate, object) triples. As an alternative, images offer a more intuitive way of understanding event semantics. We observe that event described in text and in images show different abstraction levels and therefore should be projected onto heterogeneous embedding spaces, as opposed to what have been done in previous approaches which project signals from different modalities onto a homogeneous space. In this paper, we propose a Multimodal Event Representation Learning framework (MERL) to learn event representations based on both text and image modalities simultaneously. Event textual triples are projected as Gaussian density embeddings by a dual-path Gaussian triple encoder, while event images are projected as point embeddings by a visual event component-aware image encoder. Moreover, a novel score function motivated by statistical hypothesis testing is introduced to coordinate two embedding spaces. Experiments are conducted on various multimodal event-related tasks and results show that MERL outperforms a number of unimodal and multimodal baselines, demonstrating the effectiveness of the proposed framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Multi-event Video-Text RetrievalGengyuan Zhang, Jisen Ren, Jindong Gu, Volker TrespICCV 2023 · 被引用 19 次
- Multimedia Event Extraction From News With a Unified Contrastive Learning FrameworkJian Liu, Yufeng Chen, Jinan XuACM MM 2022 · 被引用 12 次
- DyMRL: Dynamic Multispace Representation Learning for Multimodal Event Forecasting in Knowledge GraphFeng Zhao, Kangzheng Liu, Teng Peng, Yu Yang 等WWW 2026
- Image Enhanced Event Detection in News ArticlesMeihan Tong, Shuai Wang, Yixin Cao, Bin Xu 等AAAI 2020 · 被引用 43 次
- Three Stream Based Multi-level Event Contrastive Learning for Text-Video Event ExtractionJiaqi Li, Chuanyi Zhang, Miaozeng Du, Dehai Min 等EMNLP 2023 · 被引用 1 次
