MERL: Multimodal Event Representation Learning in Heterogeneous Embedding Spaces
Linhai Zhang, Deyu Zhou, Yulan He, Zeng Yang
Abstract
Previous work has shown the effectiveness of using event representations for tasks such as script event prediction and stock market prediction. It is however still challenging to learn the subtle semantic differences between events based solely on textual descriptions of events often represented as (subject, predicate, object) triples. As an alternative, images offer a more intuitive way of understanding event semantics. We observe that event described in text and in images show different abstraction levels and therefore should be projected onto heterogeneous embedding spaces, as opposed to what have been done in previous approaches which project signals from different modalities onto a homogeneous space. In this paper, we propose a Multimodal Event Representation Learning framework (MERL) to learn event representations based on both text and image modalities simultaneously. Event textual triples are projected as Gaussian density embeddings by a dual-path Gaussian triple encoder, while event images are projected as point embeddings by a visual event component-aware image encoder. Moreover, a novel score function motivated by statistical hypothesis testing is introduced to coordinate two embedding spaces. Experiments are conducted on various multimodal event-related tasks and results show that MERL outperforms a number of unimodal and multimodal baselines, demonstrating the effectiveness of the proposed framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- Multi-event Video-Text RetrievalGengyuan Zhang, Jisen Ren, Jindong Gu, Volker TrespICCV 2023 · 19 citations
- Multimedia Event Extraction From News With a Unified Contrastive Learning FrameworkJian Liu, Yufeng Chen, Jinan XuACM MM 2022 · 12 citations
- DyMRL: Dynamic Multispace Representation Learning for Multimodal Event Forecasting in Knowledge GraphFeng Zhao, Kangzheng Liu, Teng Peng, Yu Yang et al.WWW 2026
- Image Enhanced Event Detection in News ArticlesMeihan Tong, Shuai Wang, Yixin Cao, Bin Xu et al.AAAI 2020 · 43 citations
- Three Stream Based Multi-level Event Contrastive Learning for Text-Video Event ExtractionJiaqi Li, Chuanyi Zhang, Miaozeng Du, Dehai Min et al.EMNLP 2023 · 1 citation
