EvEntS ReaLM: Event Reasoning of Entity States via Language Models
Evangelia Spiliopoulou, Artidoro Pagnoni, Yonatan Bisk, Eduard H. Hovy
Abstract
This paper investigates models of event implications. Specifically, how well models predict entity state-changes, by targeting their understanding of physical attributes. Nominally, Large Language models (LLM) have been exposed to procedural knowledge about how objects interact, yet our benchmarking shows they fail to reason about the world. Conversely, we also demonstrate that existing approaches often misrepresent the surprising abilities of LLMs via improper task encodings and that proper model prompting can dramatically improve performance of reported baseline results across multiple tasks. In particular, our results indicate that our prompting technique is especially useful for unseen attributes (out-of-domain) or when only limited data is available. 1 * Equal contribution. † Work completed before joining AWS AI Labs. 1 https://github.com/spilioeve/eventsrealm Context: The robot holds a laptop. The robot forcefully throws the laptop. Context: Pick up the yogurt, bananas, and sorbet. Place the ingredients in a blender. Blend the mixture until it's smooth in texture. PiGLET Open PI What attributes changed: Laptop is broken, picked-up and its location is different. What attributes changed: 1. The cleanness, weight, volume and fullness of the blender changed. 2. The texture and appearance of the mixture changed. Query: "" Target: n-dim binary vector, n = #attributes Query each attribute in candidate list Query1: Is the location of the mug different? Target: The location of the mug is different. Query2: Is the temperature of the mug different? Target: The temperature of the mug is unchanged.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2ab0f3f-ae12-417f-9a34-5c944e389a5cCited by top-tier papers4
- Fundamental Capabilities of Large Language Models and their Applications in Domain Scenarios: A SurveyJiawei Li, Yizhe Yang, Yu Bai, Xiaofeng Zhou et al.ACL 2024 · 15 citations
- Understand the Dynamic World: An End-to-End Knowledge Informed Framework for Open Domain Entity State TrackingMingchen Li, Lifu HuangSIGIR 2023 · 6 citations
- From Heuristic to Analytic: Cognitively Motivated Strategies for Coherent Physical Commonsense ReasoningZheyuan Zhang, Shane Storks, Fengyuan Hu, Sungryull Sohn et al.EMNLP 2023
- ActionReasoningBench: Reasoning about Actions with and without Ramification ConstraintsDivij Handa, Pavel Dolin, Shrinidhi Kumbhar, Tran Cao Son et al.ICLR 2025
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
Related papers
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha et al.EMNLP 2023 · 17 citations
- Can we trust LLM Self-Explanations for Entity Resolution?Tommaso Teofili, Donatella Firmani, Nick Koudas, Paolo Merialdo et al.VLDB 2026 · 2 citations
- Language Models Can Improve Event Prediction by Few-Shot Abductive ReasoningXiaoming Shi, Siqiao Xue, Kangrui Wang, Fan Zhou et al.NeurIPS 2023 · 95 citations
- Action Boundary Blindness: When LLM Agents Cannot Tell Where One Action Ends and Another BeginsZhangyi Wang, Bingnan Yu, Jiexiang Xu, Zongze LiACL 2026
- A Comprehensive Evaluation on Event Reasoning of Large Language ModelsZhengwei Tao, Zhi Jin, Yifan Zhang, Xiancai Chen et al.AAAI 2025 · 8 citations
