Beyond Surface Simplicity: Revealing Hidden Reasoning Attributes for Precise Commonsense Diagnosis
Huijun Lian, Zekai Sun, Keqi Chen, Yingming Gao, Ya Li
Abstract
Commonsense question answering (QA) are widely used to evaluate the commonsense abilities of large language models. However, answering commonsense questions correctly requires not only knowledge but also reasoning-even for seemingly simple questions. We demonstrate that such hidden reasoning attributes in commonsense questions can lead evaluation accuracy differences of up to 24.8% across different difficulty levels in the same benchmark. Current benchmarks overlook these hidden reasoning attributes, making it difficult to assess a model's specific levels of commonsense knowledge and reasoning ability. To address this issue, we introduce Re-ComSBench, a novel framework that reveals hidden reasoning attributes behind commonsense questions by leveraging the knowledge generated during the reasoning process. Additionally, ReComSBench proposes three new metrics for decoupled evaluation: Knowledge Balanced Accuracy, Marginal Sampling Gain, and Knowledge Coverage Ratio. Experiments show that ReComSBench provides insights into model performance that traditional benchmarks cannot offer. The difficulty stratification based on revealed hidden reasoning attributes performs as effectively as the model-probabilitybased approach but is more generalizable and better suited for improving a model's commonsense reasoning abilities. By uncovering and analyzing the hidden reasoning attributes in commonsense data, ReComSBench offers a new approach to enhancing existing commonsense benchmarks. * Corresponding author Quesion Where do all animals live? Question and Options (D). meadow; (E). zoos. Variables : set of places from options. : x is an animal. : animal x can live in place p. : p is a universal habitat.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72f47ed3-c52f-4b2f-a20b-638918fdfa11Builds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
Related papers
- ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question AnsweringFrancesco Maria Molfese, Luca Moroni, Ciro Porcaro, Simone Conia et al.ACL 2026 · 1 citation
- ZEBRA: Zero-Shot Example-Based Retrieval Augmentation for Commonsense Question AnsweringFrancesco Molfese, Simone Conia, Riccardo Orlando, Roberto NavigliEMNLP 2024
- Explore What LLM Does Not Know in Complex Question AnsweringXin Lin, Zhenya Huang, Zhiqiang Zhang, Jun Zhou et al.AAAI 2025 · 8 citations
- AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World ScenariosLisa Alazraki, Lihu Chen, Ana Brassard, Joe Stacey et al.ACL 2026 · 1 citation
- MME-Reasoning: A Broad-Spectrum Benchmark for Evaluating Logical Reasoning in MLLMsJiakang Yuan, Tianshuo Peng, Yilei Jiang, Yiting Lu et al.ICML 2026
