Benchmarking Commonsense Knowledge Base Population with an Effective Evaluation Dataset
Tianqing Fang, Weiqi Wang, Sehyun Choi, Shibo Hao, Hongming Zhang, Yangqiu Song, Bin He
Abstract
Reasoning over commonsense knowledge bases (CSKBs) whose elements are in the form of free-text is an important yet hard task in NLP. While CSKB completion only fills the missing links within the domain of the CSKB, CSKB population is alternatively proposed with the goal of reasoning unseen assertions from external resources. In this task, CSKBs are grounded to a large-scale eventuality (activity, state, and event) graph to discriminate whether novel triples from the eventuality graph are plausible or not. However, existing evaluations on the population task are either not accurate (automatic evaluation with randomly sampled negative examples) or of small scale (human annotation). In this paper, we benchmark the CSKB population task with a new large-scale dataset by first aligning four popular CSKBs, and then presenting a highquality human-annotated evaluation set to probe neural models' commonsense reasoning ability. We also propose a novel inductive commonsense reasoning model that reasons over graphs. Experimental results show that generalizing commonsense reasoning on unseen assertions is inherently a hard task. Models achieving high accuracy during training perform poorly on the evaluation set, with a large gap between human performance. Codes and data are available at https://github.com/ HKUST-KnowComp/CSKB-Population .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d8f92e5-9719-42db-b591-c26843f69e75Cited by top-tier papers6
- EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product AssociationWeiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag et al.ACL 2025 · 15 citations
- CAT: A Contextualized Conceptualization and Instantiation Framework for Commonsense ReasoningWeiqi Wang, Tianqing Fang, Baixuan Xu, Chun Yi Louis Bo et al.ACL 2023 · 13 citations
- CANDLE: Iterative Conceptualization and Instantiation Distillation from Large Language Models for Commonsense ReasoningWeiqi Wang, Tianqing Fang, Chunyang Li, Haochen Shi et al.ACL 2024 · 10 citations
- Complex Reasoning over Logical Queries on Commonsense Knowledge GraphsTianqing Fang, Zeming Chen, Yangqiu Song, Antoine BosselutACL 2024 · 5 citations
- Dense-ATOMIC: Towards Densely-connected ATOMIC with High Knowledge Coverage and Massive Multi-hop PathsXiangqing Shen, Siwei Wu, Rui XiaACL 2023 · 3 citations
Builds on9
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsJena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da et al.AAAI 2021 · 458 citations
- Commonsense Knowledge Base Completion with Structural and Semantic ContextChaitanya Malaviya, Chandra Bhagavatula, Antoine Bosselut, Yejin ChoiAAAI 2020 · 155 citations
Related papers
- DISCOS: Bridging the Gap between Discourse Knowledge and Commonsense KnowledgeTianqing Fang, Hongming Zhang, Weiqi Wang, Yangqiu Song et al.WWW 2021 · 48 citations
- ACCENT: An Automatic Event Commonsense Evaluation Metric for Open-Domain Dialogue SystemsSarik Ghazarian, Yijia Shao, Rujun Han, Aram Galstyan et al.ACL 2023 · 3 citations
- A Foundation Model for Zero-shot Logical Query ReasoningMichael Galkin, Jincheng Zhou, Bruno Ribeiro, Jian Tang et al.NeurIPS 2024 · 20 citations
- CAKE: A Scalable Commonsense-Aware Framework For Multi-View Knowledge Graph CompletionGuanglin Niu, Bo Li, Yongfei Zhang, Shiliang PuACL 2022 · 56 citations
- ASER: A Large-scale Eventuality Knowledge GraphHongming Zhang, Xin Liu, Haojie Pan, Yangqiu Song et al.WWW 2020 · 183 citations
