EvEntS ReaLM: Event Reasoning of Entity States via Language Models
Evangelia Spiliopoulou, Artidoro Pagnoni, Yonatan Bisk, Eduard H. Hovy
摘要
This paper investigates models of event implications. Specifically, how well models predict entity state-changes, by targeting their understanding of physical attributes. Nominally, Large Language models (LLM) have been exposed to procedural knowledge about how objects interact, yet our benchmarking shows they fail to reason about the world. Conversely, we also demonstrate that existing approaches often misrepresent the surprising abilities of LLMs via improper task encodings and that proper model prompting can dramatically improve performance of reported baseline results across multiple tasks. In particular, our results indicate that our prompting technique is especially useful for unseen attributes (out-of-domain) or when only limited data is available. 1 * Equal contribution. † Work completed before joining AWS AI Labs. 1 https://github.com/spilioeve/eventsrealm Context: The robot holds a laptop. The robot forcefully throws the laptop. Context: Pick up the yogurt, bananas, and sorbet. Place the ingredients in a blender. Blend the mixture until it's smooth in texture. PiGLET Open PI What attributes changed: Laptop is broken, picked-up and its location is different. What attributes changed: 1. The cleanness, weight, volume and fullness of the blender changed. 2. The texture and appearance of the mixture changed. Query: "" Target: n-dim binary vector, n = #attributes Query each attribute in candidate list Query1: Is the location of the mug different? Target: The location of the mug is different. Query2: Is the temperature of the mug different? Target: The temperature of the mug is unchanged.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Fundamental Capabilities of Large Language Models and their Applications in Domain Scenarios: A SurveyJiawei Li, Yizhe Yang, Yu Bai, Xiaofeng Zhou 等ACL 2024 · 被引用 15 次
- Understand the Dynamic World: An End-to-End Knowledge Informed Framework for Open Domain Entity State TrackingMingchen Li, Lifu HuangSIGIR 2023 · 被引用 6 次
- From Heuristic to Analytic: Cognitively Motivated Strategies for Coherent Physical Commonsense ReasoningZheyuan Zhang, Shane Storks, Fengyuan Hu, Sungryull Sohn 等EMNLP 2023
- ActionReasoningBench: Reasoning about Actions with and without Ramification ConstraintsDivij Handa, Pavel Dolin, Shrinidhi Kumbhar, Tran Cao Son 等ICLR 2025
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
相关 Paper
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha 等EMNLP 2023 · 被引用 17 次
- Can we trust LLM Self-Explanations for Entity Resolution?Tommaso Teofili, Donatella Firmani, Nick Koudas, Paolo Merialdo 等VLDB 2026 · 被引用 2 次
- Language Models Can Improve Event Prediction by Few-Shot Abductive ReasoningXiaoming Shi, Siqiao Xue, Kangrui Wang, Fan Zhou 等NeurIPS 2023 · 被引用 95 次
- Action Boundary Blindness: When LLM Agents Cannot Tell Where One Action Ends and Another BeginsZhangyi Wang, Bingnan Yu, Jiexiang Xu, Zongze LiACL 2026
- A Comprehensive Evaluation on Event Reasoning of Large Language ModelsZhengwei Tao, Zhi Jin, Yifan Zhang, Xiancai Chen 等AAAI 2025 · 被引用 8 次
