Benchmarking Commonsense Knowledge Base Population with an Effective Evaluation Dataset
Tianqing Fang, Weiqi Wang, Sehyun Choi, Shibo Hao, Hongming Zhang, Yangqiu Song, Bin He
摘要
Reasoning over commonsense knowledge bases (CSKBs) whose elements are in the form of free-text is an important yet hard task in NLP. While CSKB completion only fills the missing links within the domain of the CSKB, CSKB population is alternatively proposed with the goal of reasoning unseen assertions from external resources. In this task, CSKBs are grounded to a large-scale eventuality (activity, state, and event) graph to discriminate whether novel triples from the eventuality graph are plausible or not. However, existing evaluations on the population task are either not accurate (automatic evaluation with randomly sampled negative examples) or of small scale (human annotation). In this paper, we benchmark the CSKB population task with a new large-scale dataset by first aligning four popular CSKBs, and then presenting a highquality human-annotated evaluation set to probe neural models' commonsense reasoning ability. We also propose a novel inductive commonsense reasoning model that reasons over graphs. Experimental results show that generalizing commonsense reasoning on unseen assertions is inherently a hard task. Models achieving high accuracy during training perform poorly on the evaluation set, with a large gap between human performance. Codes and data are available at https://github.com/ HKUST-KnowComp/CSKB-Population .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product AssociationWeiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag 等ACL 2025 · 被引用 15 次
- CAT: A Contextualized Conceptualization and Instantiation Framework for Commonsense ReasoningWeiqi Wang, Tianqing Fang, Baixuan Xu, Chun Yi Louis Bo 等ACL 2023 · 被引用 13 次
- CANDLE: Iterative Conceptualization and Instantiation Distillation from Large Language Models for Commonsense ReasoningWeiqi Wang, Tianqing Fang, Chunyang Li, Haochen Shi 等ACL 2024 · 被引用 10 次
- Complex Reasoning over Logical Queries on Commonsense Knowledge GraphsTianqing Fang, Zeming Chen, Yangqiu Song, Antoine BosselutACL 2024 · 被引用 5 次
- Dense-ATOMIC: Towards Densely-connected ATOMIC with High Knowledge Coverage and Massive Multi-hop PathsXiangqing Shen, Siwei Wu, Rui XiaACL 2023 · 被引用 3 次
它引用的顶会 Paper9
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsJena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da 等AAAI 2021 · 被引用 458 次
- Commonsense Knowledge Base Completion with Structural and Semantic ContextChaitanya Malaviya, Chandra Bhagavatula, Antoine Bosselut, Yejin ChoiAAAI 2020 · 被引用 155 次
相关 Paper
- DISCOS: Bridging the Gap between Discourse Knowledge and Commonsense KnowledgeTianqing Fang, Hongming Zhang, Weiqi Wang, Yangqiu Song 等WWW 2021 · 被引用 48 次
- ACCENT: An Automatic Event Commonsense Evaluation Metric for Open-Domain Dialogue SystemsSarik Ghazarian, Yijia Shao, Rujun Han, Aram Galstyan 等ACL 2023 · 被引用 3 次
- A Foundation Model for Zero-shot Logical Query ReasoningMichael Galkin, Jincheng Zhou, Bruno Ribeiro, Jian Tang 等NeurIPS 2024 · 被引用 20 次
- CAKE: A Scalable Commonsense-Aware Framework For Multi-View Knowledge Graph CompletionGuanglin Niu, Bo Li, Yongfei Zhang, Shiliang PuACL 2022 · 被引用 56 次
- ASER: A Large-scale Eventuality Knowledge GraphHongming Zhang, Xin Liu, Haojie Pan, Yangqiu Song 等WWW 2020 · 被引用 183 次
