CN-AutoMIC: Distilling Chinese Commonsense Knowledge from Pretrained Language Models
Chenhao Wang, Jiachun Li, Yubo Chen, Kang Liu, Jun Zhao
摘要
Commonsense knowledge graphs (CKGs) are increasingly applied in various natural language processing tasks. However, most existing CKGs are limited to English, which hinders related research in non-English languages. Meanwhile, directly generating commonsense knowledge from pretrained language models has recently received attention, yet it has not been explored in non-English languages. In this paper, we propose a large-scale Chinese CKG generated from multilingual PLMs, named as CN-AutoMIC, aiming to fill the research gap of non-English CKGs. To improve the efficiency, we propose generate-by-category strategy to reduce invalid generation. To ensure the filtering quality, we develop cascaded filters to discard low-quality results. To further increase the diversity and density, we introduce a bootstrapping iteration process to reuse generated results. Finally, we conduct detailed analyses on CN-AutoMIC from different aspects. Empirical results show the proposed CKG has high quality and diversity, surpassing the direct translation version of similar English CKGs. We also find some interesting deficiency patterns and differences between relations, which reveal pending problems in commonsense knowledge generation. We share the resources and related models for further study.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsJena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da 等AAAI 2021 · 被引用 458 次
- Automated Storytelling via Causal, Commonsense Plot OrderingPrithviraj Ammanabrolu, Wesley Cheung, William Broniec, Mark O. RiedlAAAI 2021 · 被引用 72 次
- DISCOS: Bridging the Gap between Discourse Knowledge and Commonsense KnowledgeTianqing Fang, Hongming Zhang, Weiqi Wang, Yangqiu Song 等WWW 2021 · 被引用 48 次
相关 Paper
- ATAP: Automatic Template-Augmented Commonsense Knowledge Graph Completion via Pre-Trained Language ModelsFu Zhang, Yifan Ding, Jingwei ChengEMNLP 2024 · 被引用 2 次
- Diverse and Informative Dialogue Generation with Context-Specific Commonsense Knowledge AwarenessSixing Wu, Ying Li, Dawei Zhang, Yang Zhou 等ACL 2020 · 被引用 104 次
- Generated Knowledge Prompting for Commonsense ReasoningJiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck 等ACL 2022
- Prix-LM: Pretraining for Multilingual Knowledge Base ConstructionWenxuan Zhou, Fangyu Liu, Ivan Vulic, Nigel Collier 等ACL 2022 · 被引用 21 次
- Enhancing Multilingual Language Model with Massive Multilingual Knowledge TriplesLinlin Liu, Xin Li, Ruidan He, Lidong Bing 等EMNLP 2022 · 被引用 15 次
