CN-AutoMIC: Distilling Chinese Commonsense Knowledge from Pretrained Language Models
Chenhao Wang, Jiachun Li, Yubo Chen, Kang Liu, Jun Zhao
Abstract
Commonsense knowledge graphs (CKGs) are increasingly applied in various natural language processing tasks. However, most existing CKGs are limited to English, which hinders related research in non-English languages. Meanwhile, directly generating commonsense knowledge from pretrained language models has recently received attention, yet it has not been explored in non-English languages. In this paper, we propose a large-scale Chinese CKG generated from multilingual PLMs, named as CN-AutoMIC, aiming to fill the research gap of non-English CKGs. To improve the efficiency, we propose generate-by-category strategy to reduce invalid generation. To ensure the filtering quality, we develop cascaded filters to discard low-quality results. To further increase the diversity and density, we introduce a bootstrapping iteration process to reuse generated results. Finally, we conduct detailed analyses on CN-AutoMIC from different aspects. Empirical results show the proposed CKG has high quality and diversity, surpassing the direct translation version of similar English CKGs. We also find some interesting deficiency patterns and differences between relations, which reveal pending problems in commonsense knowledge generation. We share the resources and related models for further study.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a16b952-9b89-4c0d-b319-bd55d7ad6a58Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsJena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da et al.AAAI 2021 · 458 citations
- Automated Storytelling via Causal, Commonsense Plot OrderingPrithviraj Ammanabrolu, Wesley Cheung, William Broniec, Mark O. RiedlAAAI 2021 · 72 citations
- DISCOS: Bridging the Gap between Discourse Knowledge and Commonsense KnowledgeTianqing Fang, Hongming Zhang, Weiqi Wang, Yangqiu Song et al.WWW 2021 · 48 citations
Related papers
- ATAP: Automatic Template-Augmented Commonsense Knowledge Graph Completion via Pre-Trained Language ModelsFu Zhang, Yifan Ding, Jingwei ChengEMNLP 2024 · 2 citations
- Diverse and Informative Dialogue Generation with Context-Specific Commonsense Knowledge AwarenessSixing Wu, Ying Li, Dawei Zhang, Yang Zhou et al.ACL 2020 · 104 citations
- Generated Knowledge Prompting for Commonsense ReasoningJiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck et al.ACL 2022
- Prix-LM: Pretraining for Multilingual Knowledge Base ConstructionWenxuan Zhou, Fangyu Liu, Ivan Vulic, Nigel Collier et al.ACL 2022 · 21 citations
- Enhancing Multilingual Language Model with Massive Multilingual Knowledge TriplesLinlin Liu, Xin Li, Ruidan He, Lidong Bing et al.EMNLP 2022 · 15 citations
