CRoW: Benchmarking Commonsense Reasoning in Real-World Tasks
Mete Ismayilzada, Debjit Paul, Syrielle Montariol, Mor Geva, Antoine Bosselut
Abstract
Recent efforts in natural language processing (NLP) commonsense reasoning research have yielded a considerable number of new datasets and benchmarks. However, most of these datasets formulate commonsense reasoning challenges in artificial scenarios that are not reflective of the tasks which real-world NLP systems are designed to solve. In this work, we present CROW, a manually-curated, multitask benchmark that evaluates the ability of models to apply commonsense reasoning in the context of six real-world NLP tasks. CROW is constructed using a multi-stage data collection pipeline that rewrites examples from existing datasets using commonsense-violating perturbations. We use CROWto study how NLP systems perform across different dimensions of commonsense knowledge, such as physical, temporal, and social reasoning. We find a significant performance gap when NLP systems are evaluated on CROWcompared to humans, showcasing that commonsense reasoning is far from being solved in real-world task settings. We make our dataset and leaderboard available to the research community. 1 * Equal contribution 1 https://github.com/mismayil/crow Dialogue Agent: Hi, would you like some free candies? Human: Sure. What are you handing these out for? Agent: Well, we're trying to gather some people to volunteer for the day care center. Human: Uh… Agent: It's OK. You don't have to volunteer if you eat the candies. Agent: It's OK. You don't have to eat the candies if you volunteer. Artificial Evaluation CRoW Evaluation (real-world) doesn't have prerequisite NOT OK OK OK Bob decided to volunteer because he wanted to eat the candies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- ATLAS: Constraints-Aware Multi-Agent Collaboration for Real-World Travel PlanningJihye Choi, Jinsung Yoon, Jiefeng Chen, Somesh Jha et al.ICLR 2026 · 18 citations
- Rule or Story, Which is a Better Commonsense Expression for Talking with Large Language Models?Ning Bian, Xianpei Han, Hongyu Lin, Yaojie Lu et al.ACL 2024 · 1 citation
- EmoBench: Evaluating the Emotional Intelligence of Large Language ModelsSahand Sabour, Siyang Liu, Zheyuan Zhang, June M. Liu et al.ACL 2024
- Connecting the Knowledge Dots: Retrieval-augmented Knowledge Connection for Commonsense ReasoningJunho Kim, Soyeon Bak, Mingyu Lee, Minju Hong et al.EMNLP 2025
Builds on15
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi et al.ICLR 2020 · 521 citations
- (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsJena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da et al.AAAI 2021 · 458 citations
Related papers
- SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World KnowledgeAndong Wang, Bo Wu, Sunli Chen, Zhenfang Chen et al.CVPR 2024 · 7 citations
- Common Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense ReasoningBill Yuchen Lin, Seyeon Lee, Xiaoyang Qiao, Xiang RenACL 2021
- CORECODE: A Common Sense Annotated Dialogue Dataset with Benchmark Tasks for Chinese Large Language ModelsDan Shi, Chaobin You, Jiantao Huang, Taihao Li et al.AAAI 2024 · 3 citations
- SpaCE-Eval: A Benchmark for Real-World Multi-Modal ReasoningXuyou Yang, Yucheng Zhao, Wenxuan Zhang, Immanuel KohICLR 2026
- SVBench: Evaluation of Video Generation Models on Social ReasoningWenshuo Peng, Gongxuan Wang, Tianmeng Yang, Chuanhao Li et al.CVPR 2026 · 5 citations
