Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning Extraction
Wenxuan Liu, Zixuan Li, Long Bai, Yuxin Zuo, Daozhu Xu, Xiaolong Jin, Jiafeng Guo, Xueqi Cheng
摘要
Developing a general-purpose system that can extract events with massive types is a longstanding target in Event Extraction (EE). In doing so, the basic challenge comes from the absence of an efficient and effective annotation framework to construct the corresponding datasets. In this paper, we propose an LLM-based collaborative annotation framework. Through collaboration among multiple LLMs and a subsequent voting process, it refines annotations of triggers from distant supervision and then carries out argument annotation. Finally, we create EEMT, the largest EE dataset to date, featuring over 200,000 samples, 3,465 event types, and 6,297 role types. Evaluation on the human-annotated test set demonstrates that the proposed framework achieves the F1 scores of 90.1% and 85.3% for event detection and argument extraction, strongly validating its effectiveness. Besides, to alleviate the excessively long prompts caused by massive types, we propose an LLM-based Partitioning method for EE called LLM-PEE. It first recalls candidate event types and then splits them into multiple partitions for LLMs to extract. After fine-tuning on the EEMT training set, the distilled LLM-PEE with 7B parameters outperforms state-of-the-art methods by 5.4% and 6.1% in event detection and argument extraction. Besides, it also surpasses mainstream LLMs by 12.9% on the unseen datasets, which strongly demonstrates the event diversity of the EEMT dataset and the generalization capabilities of the LLM-PEE method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Event Extraction by Answering (Almost) Natural QuestionsXinya Du, Claire CardieEMNLP 2020 · 被引用 391 次
- Event Extraction as Machine Reading ComprehensionJian Liu, Yubo Chen, Kang Liu, Wei Bi 等EMNLP 2020 · 被引用 300 次
- CASIE: Extracting Cybersecurity Event Information from TextTaneeya Satyapanich, Francis Ferraro, Tim FininAAAI 2020 · 被引用 148 次
相关 Paper
- MEE: A Novel Multilingual Event Extraction DatasetAmir Pouran Ben Veyseh, Javid Ebrahimi, Franck Dernoncourt, Thien Huu NguyenEMNLP 2022 · 被引用 3 次
- MAVEN-ARG: Completing the Puzzle of All-in-One Event Understanding Dataset with Event Argument AnnotationXiaozhi Wang, Hao Peng, Yong Guan, Kaisheng Zeng 等ACL 2024
- Cross-modal Multi-task Learning for Multimedia Event ExtractionJianwei Cao, Yanli Hu, Zhen Tan, Xiang ZhaoAAAI 2025 · 被引用 8 次
- Reflective Agreement: Combining Self-Mixture of Agents with a Sequence Tagger for Robust Event ExtractionFatemeh Haji, Mazal Bethany, Cho-Yu Jason Chiang, Anthony Rios 等EMNLP 2025
- Prompt for Extraction? PAIE: Prompting Argument Interaction for Event Argument ExtractionYubo Ma, Zehao Wang, Yixin Cao, Mukai Li 等ACL 2022 · 被引用 182 次
