Cross-media Structured Common Space for Multimedia Event Extraction
Manling Li, Alireza Zareian, Qi Zeng, Spencer Whitehead, Di Lu, Heng Ji, Shih-Fu Chang
摘要
We introduce a new task, MultiMedia Event Extraction (M 2 E 2 ), which aims to extract events and their arguments from multimedia documents. We develop the first benchmark and collect a dataset of 245 multimedia news articles with extensively annotated events and arguments. 1 We propose a novel method, Weakly Aligned Structured Embedding (WASE), that encodes structured representations of semantic information from textual and visual data into a common embedding space. The structures are aligned across modalities by employing a weakly supervised training strategy, which enables exploiting available resources without explicit cross-media annotation. Compared to unimodal state-of-the-art methods, our approach achieves 4.0% and 9.8% absolute F-score gains on text event argument role labeling and visual event extraction. Compared to stateof-the-art multimedia unstructured representations, we achieve 8.3% and 5.0% absolute Fscore gains on multimedia event extraction and argument role labeling, respectively. By utilizing images, we extract 21.4% more event mentions than traditional text-only methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- CLIP-Event: Connecting Text and Images with Event StructuresManling Li, Ruochen Xu, Shuohang Wang, Luowei Zhou 等CVPR 2022 · 被引用 103 次
- Language Models Can Improve Event Prediction by Few-Shot Abductive ReasoningXiaoming Shi, Siqiao Xue, Kangrui Wang, Fan Zhou 等NeurIPS 2023 · 被引用 95 次
- Text2Mol: Cross-Modal Molecule Retrieval with Natural Language QueriesCarl Edwards, ChengXiang Zhai, Heng JiEMNLP 2021 · 被引用 79 次
- MuMuQA: Multimedia Multi-Hop News Question Answering via Cross-Media Knowledge Extraction and GroundingRevanth Gangi Reddy, Xilin Rui, Manling Li, Xudong Lin 等AAAI 2022 · 被引用 37 次
- Timeline Summarization based on Event Graph Compression via Time-Aware Optimal TransportManling Li, Tengfei Ma, Mo Yu, Lingfei Wu 等EMNLP 2021 · 被引用 25 次
它引用的顶会 Paper5
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal RetrievalShupeng Su, Zhisheng Zhong, Chao ZhangICCV 2019 · 被引用 261 次
- Adversarial Representation Learning for Text-to-Image MatchingNikolaos Sarafianos, Xiang Xu, Ioannis A. KakadiarisICCV 2019 · 被引用 228 次
相关 Paper
- Cross-modal Multi-task Learning for Multimedia Event ExtractionJianwei Cao, Yanli Hu, Zhen Tan, Xiang ZhaoAAAI 2025 · 被引用 8 次
- Multimedia Event Extraction From News With a Unified Contrastive Learning FrameworkJian Liu, Yufeng Chen, Jinan XuACM MM 2022 · 被引用 12 次
- Multimedia Event Extraction with LLM Knowledge EditingJiaao Yu, Yijing Lin, Zhipeng Gao, Xuesong Qiu 等EMNLP 2025
- Training Multimedia Event Extraction With Generated Images and CaptionsZilin Du, Yunxin Li, Xu Guo, Yidan Sun 等ACM MM 2023 · 被引用 7 次
- Beyond Grounding: Extracting Fine-Grained Event Hierarchies across ModalitiesHammad A. Ayyubi, Christopher Thomas, Lovish Chum, Rahul Lokesh 等AAAI 2024 · 被引用 2 次
