Multi-View Clustering for Open Knowledge Base Canonicalization
Wei Shen, Yang Yang, Yinan Liu
Abstract
Open information extraction (OIE) methods extract plenty of OIE triples <noun phrase, relation phrase, noun phrase> from unstructured text, which compose large open knowledge bases (OKBs). Noun phrases and relation phrases in such OKBs are not canonicalized, which leads to scattered and redundant facts. It is found that two views of knowledge (i.e., a fact view based on the fact triple and a context view based on the fact triple's source context) provide complementary information that is vital to the task of OKB canonicalization, which clusters synonymous noun phrases and relation phrases into the same group and assigns them unique identifiers. However, these two views of knowledge have so far been leveraged in isolation by existing works. In this paper, we propose CMVC, a novel unsupervised framework that leverages these two views of knowledge jointly for canonicalizing OKBs without the need of manually annotated labels. To achieve this goal, we pro- pose a multi-view CH K-Means clustering algorithm to mutually reinforce the clustering of view-specific embeddings learned from each view by considering their different clustering qualities. In order to further enhance the canonicalization performance, we propose a training data optimization strategy in terms of data quantity and data quality respectively in each particular view to refine the learned view-specific embeddings in an iterative manner. Additionally, we propose a Log-Jump algorithm to predict the optimal number of clusters in a data-driven way without requiring any labels. We demonstrate the superiority of our framework through extensive experiments on multiple real-world OKB data sets against state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0aaaca74-bf56-4027-a1fe-89349a00b56cCited by top-tier papers2
- KGE Calibrator: An Efficient Probability Calibration Method of Knowledge Graph Embedding Models for Trustworthy Link PredictionYang Yang, Mohan Timilsina, Edward CurryEMNLP 2025
- A Novel Retrieve-Read-Group Paradigm for Open Knowledge Base CanonicalizationBinhan Yang, Wei Shen, Han TianAAAI 2026
Builds on2
Related papers
- Jointly Canonicalizing and Linking Open Knowledge Base via Unified Embedding LearningWei Shen, Binhan Yang, Yinan LiuWWW 2024
- Can We Predict New Facts with Open Knowledge Graph Embeddings? A Benchmark for Open Link PredictionSamuel Broscheit, Kiril Gashteovski, Yanjie Wang, Rainer GemullaACL 2020 · 27 citations
- Knowledge Base Completion Meets Transfer LearningVid Kocijan, Thomas LukasiewiczEMNLP 2021
- Cross-view Topology Based Consistent and Complementary Information for Deep Multi-view ClusteringZhibin Dong, Siwei Wang, Jiaqi Jin, Xinwang Liu et al.ICCV 2023 · 33 citations
- When Phrases Meet Probabilities: Enabling Open Relation Extraction with Cooperating Large Language ModelsJiaxin Wang, Lingling Zhang, Wee Sun Lee, Yujie Zhong et al.ACL 2024
