Wiki2Prop: A Multimodal Approach for Predicting Wikidata Properties from Wikipedia
Michael Luggen, Julien Audiffren, Djellel Eddine Difallah, Philippe Cudré-Mauroux
Abstract
Wikidata is rapidly emerging as a key resource for a multitude of online tasks such as Speech Recognition, Entity Linking, Question Answering, or Semantic Search. The value of Wikidata is directly linked to the rich information associated with each entity – that is, the properties describing each entity as well as the relationships to other entities. Despite the tremendous manual and automatic efforts the community invested in the Wikidata project, the growing number of entities (now more than 100 million) presents multiple challenges in terms of knowledge gaps in the graph that are hard to track. To help guide the community in filling the gaps in Wikidata, we propose to identify and rank the properties that an entity might be missing. In this work, we focus on entities with a dedicated Wikipedia page in any language to make predictions directly based on textual content. We show that this problem can be formulated as a multi-label classification problem where every property defined in Wikidata is a potential label. Our main contribution, Wiki2Prop, solves this problem using a multimodal Deep Learning method to predict which properties should be attached to a given entity, using its Wikipedia page embeddings. Moreover, Wiki2Prop is able to incorporate additional features in the form of multilingual embeddings and multimodal data such as images whenever available. We empirically evaluate our approach against the state of the art and show how Wiki2Prop significantly outperforms its competitors for the task of property prediction in Wikidata, and how the use of multilingual and multimodal data improves the results further. Finally, we make Wiki2Prop available as a property recommender system that can be activated and used directly in the context of a Wikidata entity page.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62eb0ee9-35f1-4502-aa5e-a7959bb76940Cited by top-tier papers2
- MMKGR: Multi-hop Multi-modal Knowledge Graph ReasoningShangfei Zheng, Weiqing Wang, Jianfeng Qu, Hongzhi Yin et al.ICDE 2023 · 40 citations
- CAMul: Calibrated and Accurate Multi-view Time-Series ForecastingHarshavardhan Kamarthi, Lingkai Kong, Alexander Rodríguez, Chao Zhang et al.WWW 2022 · 23 citations
Related papers
- Increasing Coverage and Precision of Textual Information in Multilingual Knowledge GraphsSimone Conia, Min Li, Daniel Lee, Umar Farooq Minhas et al.EMNLP 2023 · 3 citations
- Multi-level Mixture of Experts for Multimodal Entity LinkingZhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li et al.KDD 2025 · 1 citation
- Wikidata as a seed for Web ExtractionKunpeng Guo, Dennis Diefenbach, Antoine Gourru, Christophe GravierWWW 2023 · 6 citations
- A Dual-Way Enhanced Framework from Text Matching Point of View for Multimodal Entity LinkingShezheng Song, Shan Zhao, Chengyu Wang, Tianwei Yan et al.AAAI 2024
- Cross-Lingual Phrase RetrievalHeqi Zheng, Xiao Zhang, Zewen Chi, Heyan Huang et al.ACL 2022 · 1 citation
