MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset
Weiqi Wang, Yangqiu Song
Abstract
To enable Large Language Models (LLMs) to function as conscious agents with generalizable reasoning capabilities, it is crucial that they possess the ability to comprehend situational changes (transitions) in distribution triggered by environmental factors or actions from other agents. Despite its fundamental significance, this ability remains underexplored due to the complexity of modeling infinite possible changes in an event and their associated distributions, coupled with the lack of benchmark data with situational transitions. Addressing these gaps, we propose a novel formulation of reasoning with distributional changes as a three-step discriminative process, termed as MetAphysical ReaSoning. We then introduce the first-ever benchmark, MARS, comprising three tasks corresponding to each step. These tasks systematically assess LLMs' capabilities in reasoning the plausibility of (i) changes in actions, (ii) states caused by changed actions, and (iii) situational transitions driven by changes in action. Extensive evaluations with 20 (L)LMs of varying sizes and methods indicate that all three tasks in this process pose significant challenges, even after fine-tuning. Further analyses reveal potential causes for the underperformance of LLMs and demonstrate that pre-training on largescale conceptualization taxonomies can potentially enhance LMs' metaphysical reasoning capabilities. Our data and models are publicly accessible at https://github.com/HKUST-KnowComp/MARS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8b94fa1-27cb-4e9e-b4fb-e5db4ccea469Cited by top-tier papers6
- EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product AssociationWeiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag et al.ACL 2025 · 15 citations
- CANDLE: Iterative Conceptualization and Instantiation Distillation from Large Language Models for Commonsense ReasoningWeiqi Wang, Tianqing Fang, Chunyang Li, Haochen Shi et al.ACL 2024 · 10 citations
- arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table GenerationWeiqi Wang, Jiefu Ou, Yangqiu Song, Benjamin Van Durme et al.ACL 2026 · 8 citations
- GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity TheoryWei Fan, Haoran Li, Zheye Deng, Weiqi Wang et al.EMNLP 2024 · 7 citations
- MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase UnderstandingBaixuan Xu, Weiqi Wang, Haochen Shi, Wenxuan Ding et al.EMNLP 2024 · 4 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- MME-Reasoning: A Broad-Spectrum Benchmark for Evaluating Logical Reasoning in MLLMsJiakang Yuan, Tianshuo Peng, Yilei Jiang, Yiting Lu et al.ICML 2026
- Exploring the Capacity of Pretrained Language Models for Reasoning about Actions and ChangeWeinan He, Canming Huang, Zhanhao Xiao, Yongmei LiuACL 2023
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent SystemsXuanming Zhang, Yuxuan Chen, Samuel (Min-Hsuan) Yeh, Sharon LiNeurIPS 2025 · 14 citations
- AwarenessBench: Assessing Cognitive Capabilities of Language ModelsXiaojian Li, Rongwu Xu, Tianyun Zhang, Yue Wang et al.ACL 2026
- Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive ReasoningYian Li, Wentao Tian, Yang Jiao, Tianwen Qian et al.ACM MM 2025 · 16 citations
