MAVEN-ARG: Completing the Puzzle of All-in-One Event Understanding Dataset with Event Argument Annotation
Xiaozhi Wang, Hao Peng, Yong Guan, Kaisheng Zeng, Jianhui Chen, Lei Hou, Xu Han, Yankai Lin, Zhiyuan Liu, Ruobing Xie, Jie Zhou, Juanzi Li
Abstract
Understanding events in texts is a core objective of natural language understanding, which requires detecting event occurrences, extracting event arguments, and analyzing inter-event relationships. However, due to the annotation challenges brought by task complexity, a largescale dataset covering the full process of event understanding has long been absent. In this paper, we introduce MAVEN-ARG, which augments MAVEN datasets with event argument annotations, making the first all-in-one dataset supporting event detection, event argument extraction (EAE), and event relation extraction. As an EAE benchmark, MAVEN-ARG offers three main advantages: (1) a comprehensive schema covering 162 event types and 612 argument roles, all with expert-written definitions and examples; (2) a large data scale, containing 98, 591 events and 290, 613 arguments obtained with laborious human annotation; (3) the exhaustive annotation supporting all task variants of EAE, which annotates both entity and non-entity event arguments in document level. Experiments indicate that MAVEN-ARG is quite challenging for both fine-tuned EAE models and proprietary large language models (LLMs). Furthermore, to demonstrate the benefits of an all-in-one dataset, we preliminarily explore a potential application, future event prediction, with LLMs. MAVEN-ARG and our baseline codes will be publicly released. et al., 2021; Peng et al., 2023b): event detection 043 (ED), which detects event occurrences by identi-044 fying event triggers and classifying event types; 045 event argument extraction (EAE), which extracts 046 event arguments and classifies their argument roles; 047 event relation extraction (ERE), which analyzes 048 the coreference, temporal, causal, and hierarchical 049 relationships among events. 050 Despite the importance of event understand-051 ing, a large-scale dataset covering all the event 052 understanding tasks has long been absent. Es-053 tablished sentence-level event extraction (ED and 054 EAE) datasets like ACE 2005 (Walker et al., 2006) 055 and TAC KBP (Ellis et al., 2015, 2016; Getman 056 et al., 2017) do not involve event relation types 057 besides the basic coreferences. RAMS (Ebner 058 et al., 2020) and WikiEvents (Li et al., 2021) ex-059 tend EAE to the document level but do not in-060 volve event relations. ERE datasets are mostly 061 developed independently for coreference (Cybul-062
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Event Extraction by Answering (Almost) Natural QuestionsXinya Du, Claire CardieEMNLP 2020 · 391 citations
- A Joint Neural Model for Information Extraction with Global FeaturesYing Lin, Heng Ji, Fei Huang, Lingfei WuACL 2020 · 376 citations
- Event Extraction as Machine Reading ComprehensionJian Liu, Yubo Chen, Kang Liu, Wei Bi et al.EMNLP 2020 · 300 citations
- Prompt for Extraction? PAIE: Prompting Argument Interaction for Event Argument ExtractionYubo Ma, Zehao Wang, Yixin Cao, Mukai Li et al.ACL 2022 · 182 citations
- MAVEN: A Massive General Domain Event Detection DatasetXiaozhi Wang, Ziqi Wang, Xu Han, Wangyi Jiang et al.EMNLP 2020 · 143 citations
Related papers
- MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation ExtractionXiaozhi Wang, Yulin Chen, Ning Ding, Hao Peng et al.EMNLP 2022 · 35 citations
- MEE: A Novel Multilingual Event Extraction DatasetAmir Pouran Ben Veyseh, Javid Ebrahimi, Franck Dernoncourt, Thien Huu NguyenEMNLP 2022 · 3 citations
- GENEVA: Benchmarking Generalizability for Event Argument Extraction with Hundreds of Event Types and Argument RolesTanmay Parekh, I-Hung Hsu, Kuan-Hao Huang, Kai-Wei Chang et al.ACL 2023 · 5 citations
- Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning ExtractionWenxuan Liu, Zixuan Li, Long Bai, Yuxin Zuo et al.EMNLP 2025
- VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in VideosBaoyu Liang, Qile Su, Shoutai Zhu, Yuchen Liang et al.AAAI 2025 · 5 citations
