Com² : A Causal-Guided Benchmark for Exploring Complex Commonsense Reasoning in Large Language Models
Kai Xiong, Xiao Ding, Yixin Cao, Yuxiong Yan, Li Du, Yufei Zhang, Jinglong Gao, Jiaqian Liu, Bing Qin, Ting Liu
摘要
Large language models (LLMs) have mastered abundant simple and explicit commonsense knowledge through pre-training, enabling them to achieve human-like performance in simple commonsense reasoning. Nevertheless, LLMs struggle to reason with complex and implicit commonsense knowledge that is derived from simple ones (such as understanding the longterm effects of certain events), an aspect humans tend to focus on more. Existing works focus on complex tasks like math and code, while complex commonsense reasoning remains underexplored due to its uncertainty and lack of structure. To fill this gap and align with realworld concerns, we propose a benchmark Com 2 focusing on complex commonsense reasoning. We first incorporate causal event graphs to serve as structured complex commonsense. Then we adopt causal theory (e.g., intervention) to modify the causal event graphs and obtain different scenarios that meet human concerns. Finally, an LLM is employed to synthesize examples with slow thinking, which is guided by the logical relationships in the modified causal graphs. Furthermore, we use detective stories to construct a more challenging subset. Experiments show that LLMs struggle in reasoning depth and breadth, while post-training and slow thinking can alleviate this. The code and data are available at Com 2 . Question: If human ••• ••• Options: A) ••• ••• B) Flooding C) ••• ••• Slow Thinking: ••• ••• Answer: D) Longer War Question: If human ••• ••• Options: A) ••• ••• B) Flooding C) ••• ••• Slow Thinking: ••• ••• Answer: E) Breakdown Question: If human ••• ••• Options: A) ••• ••• B) Flooding C) ••• ••• Slow Thinking: ••• ••• Answer: C) Better Sleep Question: If human ••• ••• Options: A) ••• ••• B) Flooding C) ••• ••• Slow Thinking: ••• ••• Answer(s): A) New Energy B) ••• ••• A C D E F G H I Pressure ••• Success •••
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
相关 Paper
- Complex Reasoning over Logical Queries on Commonsense Knowledge GraphsTianqing Fang, Zeming Chen, Yangqiu Song, Antoine BosselutACL 2024 · 被引用 5 次
- Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?Haoang Chi, He Li, Wenjing Yang, Feng Liu 等NeurIPS 2024 · 被引用 124 次
- Can Large Language Models Infer Causation from Correlation?Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff 等ICLR 2024 · 被引用 186 次
- COLD: Causal reasOning in cLosed Daily activitiesAbhinav Joshi, Areeb Ahmad, Ashutosh ModiNeurIPS 2024 · 被引用 11 次
- Language Models of Code are Few-Shot Commonsense LearnersAman Madaan, Shuyan Zhou, Uri Alon, Yiming Yang 等EMNLP 2022 · 被引用 103 次
