A Comprehensive Evaluation on Event Reasoning of Large Language Models
Zhengwei Tao, Zhi Jin, Yifan Zhang, Xiancai Chen, Haiyan Zhao, Jia Li, Bin Liang, Chongyang Tao, Qun Liu, Kam-Fai Wong
Abstract
Event reasoning is a fundamental ability that underlies many applications. It requires event schema knowledge to perform global reasoning and needs to deal with the diversity of the interevent relations and the reasoning paradigms. How well LLMs accomplish event reasoning on various relations and reasoning paradigms remains unknown. To mitigate this disparity, we comprehensively evaluate the abilities of event reasoning of LLMs. We introduce a novel benchmark EV 2 for EValuation of EVent reasoning. EV 2 consists of two levels of evaluation of schema and instance and is comprehensive in relations and reasoning paradigms. We conduct extensive experiments on EV 2 . We find that LLMs have abilities to accomplish event reasoning but their performances are far from satisfactory. We also notice the imbalance of event reasoning abilities in LLMs. Besides, LLMs have event schema knowledge, however, they're not aligned with humans on how to utilize the knowledge. Based on these findings, we guide the LLMs in utilizing the event schema knowledge as memory leading to improvements on event reasoning. Code and Dataset are available on https://github.com/TZWwww/EV2 . Enjoy Life John decided to spend his Rebirth Day with his friends from various countries. Celebrate John invited guests to his home. His friends sang and danced together to their heart's content, telling stories under the string lights.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1020cd4-dd90-4b96-a31b-74d61772da65Cited by top-tier papers4
- Graph is a Substrate Across Data ModalitiesZiming Li, Xiao-Ming Wu, Zehong Wang, Jiazheng Li et al.ICML 2026 · 16 citations
- PIPER: Benchmarking and Prompting Event Reasoning Boundary of LLMs via Debiasing-Distillation Enhanced TuningZhicong Lu, Changyuan Tian, PeiguangLi PeiguangLi, Li Jin et al.ACL 2025 · 4 citations
- Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive ReasoningBowen Zhang, Jun Ma, Fuqiang Niu, Li Dong et al.AAAI 2026 · 1 citation
- Schema-Guided Event Reasoning: A Plug-and-Play Event Reasoning Framework Based on Large Language ModelsYuying Liu, Xuechen Zhao, Yanyi Huang, Ye Wang et al.AAAI 2026
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- ASER: A Large-scale Eventuality Knowledge GraphHongming Zhang, Xin Liu, Haojie Pan, Yangqiu Song et al.WWW 2020 · 183 citations
- Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language ModelsDaman Arora, Himanshu Gaurav Singh, MausamEMNLP 2023 · 36 citations
Related papers
- MME-Reasoning: A Broad-Spectrum Benchmark for Evaluating Logical Reasoning in MLLMsJiakang Yuan, Tianshuo Peng, Yilei Jiang, Yiting Lu et al.ICML 2026
- Improving Large Language Models in Event Relation Logical PredictionMeiqi Chen, Yubo Ma, Kaitao Song, Yixin Cao et al.ACL 2024 · 7 citations
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha et al.EMNLP 2023 · 17 citations
- A²RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark GenerationQingchuan Ma, Yuexiao Ma, Yongkang Xie, Tianyu Xie et al.ICML 2026 · 1 citation
- LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language ModelsMihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura et al.ACL 2024
