Evaluation Pitfalls and Challenges in Multimedia Event Extraction
Philipp Seeberger, Steffen Freisinger, Tobias Bocklet, Korbinian Riedhammer
摘要
Multimedia event extraction aims to jointly identify events and their arguments across multiple modalities, such as text and images, to support more comprehensive event understanding. While recent work reports steady and substantial progress, the reliability and comparability of these results critically depend on consistent and rigorous evaluation. In this work, we present the first systematic analysis of evaluation pitfalls in multimedia event extraction and identify three major sources of issues: inconsistent data processing, inconsistent task assumptions, and overly relaxed evaluation settings. We demonstrate, through a series of controlled experiments under a strict evaluation framework, that minor evaluation choices can cause large performance variations and lead to overestimation of a model's ability to ground real-world events across modalities. Our findings highlight the need for comparable evaluation standards and encourage a shift toward more rigorous evaluation in multimedia event extraction. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- MAVEN: A Massive General Domain Event Detection DatasetXiaozhi Wang, Ziqi Wang, Xu Han, Wangyi Jiang 等EMNLP 2020 · 被引用 143 次
- CLIP-Event: Connecting Text and Images with Event StructuresManling Li, Ruochen Xu, Shuohang Wang, Luowei Zhou 等CVPR 2022 · 被引用 103 次
- Cross-media Structured Common Space for Multimedia Event ExtractionManling Li, Alireza Zareian, Qi Zeng, Spencer Whitehead 等ACL 2020 · 被引用 87 次
- Image Enhanced Event Detection in News ArticlesMeihan Tong, Shuai Wang, Yixin Cao, Bin Xu 等AAAI 2020 · 被引用 43 次
相关 Paper
- Cross-modal Multi-task Learning for Multimedia Event ExtractionJianwei Cao, Yanli Hu, Zhen Tan, Xiang ZhaoAAAI 2025 · 被引用 8 次
- Multimedia Event Extraction with LLM Knowledge EditingJiaao Yu, Yijing Lin, Zhipeng Gao, Xuesong Qiu 等EMNLP 2025
- LLaVA-MS-PIT: Multi-Modal Schema-Guided Progressive Instruction Tuning for Multi-Modal Event ExtractionHui Zhang, Po Hu, Wei Emma ZhangAAAI 2026
- Beyond Grounding: Extracting Fine-Grained Event Hierarchies across ModalitiesHammad A. Ayyubi, Christopher Thomas, Lovish Chum, Rahul Lokesh 等AAAI 2024 · 被引用 2 次
- Multimedia Event Extraction From News With a Unified Contrastive Learning FrameworkJian Liu, Yufeng Chen, Jinan XuACM MM 2022 · 被引用 12 次
