CLOMO: Counterfactual Logical Modification with Large Language Models
Yinya Huang, Ruixin Hong, Hongming Zhang, Wei Shao, Zhicheng Yang, Dong Yu, Changshui Zhang, Xiaodan Liang, Linqi Song
摘要
In this study, we delve into the realm of counterfactual reasoning capabilities of large language models (LLMs). Our primary objective is to cultivate the counterfactual thought processes within LLMs and rigorously assess these processes for their validity. Specifically, we introduce a novel task, Counterfactual Logical Modification (CLOMO), and a high-quality human-annotated benchmark. In this task, LLMs must adeptly alter a given argumentative text to uphold a predetermined logical relationship. To effectively evaluate a generation model's counterfactual capabilities, we propose an innovative evaluation metric, the decomposed Self-Evaluation Score (SES) to directly evaluate the natural language output of LLMs instead of modeling the task as a multiplechoice problem. Analysis shows that the proposed automatic metric aligns well with human preference. Our experimental results show that while LLMs demonstrate a notable capacity for logical counterfactual thinking, there remains a discernible gap between their current abilities and human performance. Code and data are available at https://github. com/Eleanor-H/CLOMO .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Executable Counterfactuals: Improving LLMs' Causal Reasoning Through CodeAniket Vashishtha, Qirun Dai, Hongyuan Mei, Amit Sharma 等ICLR 2026 · 被引用 11 次
- On the Eligibility of LLMs for Counterfactual Reasoning: A Decompositional StudyShuai Yang, Qi Yang, Luoxi Tang, Yuqiao Meng 等ICLR 2026 · 被引用 9 次
它引用的顶会 Paper18
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi 等ICLR 2020 · 被引用 521 次
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 被引用 325 次
- Deep Bidirectional Language-Knowledge Graph PretrainingMichihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang 等NeurIPS 2022 · 被引用 294 次
- KG-BART: Knowledge Graph-Augmented BART for Generative Commonsense ReasoningYe Liu, Yao Wan, Lifang He, Hao Peng 等AAAI 2021 · 被引用 220 次
相关 Paper
- Towards Unifying Evaluation of Counterfactual Explanations: Leveraging Large Language Models for Human-Centric AssessmentsMarharyta Domnich, Julius Välja, Rasmus Moorits Veski, Giacomo Magnifico 等AAAI 2025 · 被引用 11 次
- Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsYanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao 等ICML 2024 · 被引用 90 次
- RecToM: A Benchmark for Evaluating Machine Theory of Mind in LLM-based Conversational Recommender SystemsMengfan Li, Xuanhua Shi, Yang DengAAAI 2026
- Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning TasksRushang Karia, Daniel Bramblett, Daksh Dobhal, Siddharth SrivastavaICLR 2025
- Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy CompetitionKehua Feng, Keyan Ding, Hongzhi Tan, Kede Ma 等ACL 2025
