CORECODE: A Common Sense Annotated Dialogue Dataset with Benchmark Tasks for Chinese Large Language Models
Dan Shi, Chaobin You, Jiantao Huang, Taihao Li, Deyi Xiong
Abstract
As an indispensable ingredient of intelligence, commonsense reasoning is crucial for large language models (LLMs) in real-world scenarios. In this paper, we propose CORE-CODE, a dataset that contains abundant commonsense knowledge manually annotated on dyadic dialogues, to evaluate the commonsense reasoning and commonsense conflict detection capabilities of Chinese LLMs. We categorize commonsense knowledge in everyday conversations into three dimensions: entity, event, and social interaction. For easy and consistent annotation, we standardize the form of commonsense knowledge annotation in open-domain dialogues as "domain: slot = value". A total of 9 domains and 37 slots are defined to capture diverse commonsense knowledge. With these pre-defined domains and slots, we collect 76,787 commonsense knowledge annotations from 19,700 dialogues through crowdsourcing. To evaluate and enhance the commonsense reasoning capability for LLMs on the curated dataset, we establish a series of dialogue-level reasoning and detection tasks, including commonsense knowledge filling, commonsense knowledge generation, commonsense conflict phrase detection, domain identification, slot identification, and event causal inference. A wide variety of existing open-source Chinese LLMs are evaluated with these tasks on our dataset. Experimental results demonstrate that these models are not competent to predict CORECODE's plentiful reasoning content, and even ChatGPT could only achieve 0.275 and 0.084 accuracy on the domain identification and slot identification tasks under the zero-shot setting. We release the data and codes of CORECODE at https://github.com/danshi777/CORECODE to promote commonsense reasoning evaluation and study of LLMs in the context of daily conversations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6723fa9-b5ed-435b-8f7c-0e97971bf116Cited by top-tier papers3
- IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware NeuronsDan Shi, Renren Jin, Tianhao Shen, Weilong Dong et al.NeurIPS 2024 · 44 citations
- Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language ModelsDan Shi, Zhuowen Han, Simon Ostermann, Renren Jin et al.ACL 2026 · 1 citation
- Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization CorrelationsJiaxing Sun, Weiquan Huang, Jiang Wu, Chenya Gu et al.ACL 2024
Builds on14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
- (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsJena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da et al.AAAI 2021 · 458 citations
- QASC: A Dataset for Question Answering via Sentence CompositionTushar Khot, Peter Clark, Michal Guerquin, Peter Jansen et al.AAAI 2020 · 387 citations
Related papers
- Language Models of Code are Few-Shot Commonsense LearnersAman Madaan, Shuyan Zhou, Uri Alon, Yiming Yang et al.EMNLP 2022 · 103 citations
- TIMEDIAL: Temporal Commonsense Reasoning in DialogLianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He et al.ACL 2021
- CDConv: A Benchmark for Contradiction Detection in Chinese ConversationsChujie Zheng, Jinfeng Zhou, Yinhe Zheng, Libiao Peng et al.EMNLP 2022 · 6 citations
- SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World KnowledgeAndong Wang, Bo Wu, Sunli Chen, Zhenfang Chen et al.CVPR 2024 · 7 citations
- PCoKG: Personality-aware Commonsense Reasoning with DebateWeijie Li, Zhongqing Wang, Guodong ZhouAAAI 2026
