Prix-LM: Pretraining for Multilingual Knowledge Base Construction
Wenxuan Zhou, Fangyu Liu, Ivan Vulic, Nigel Collier, Muhao Chen
Abstract
Knowledge bases (KBs) contain plenty of structured world and commonsense knowledge. As such, they often complement distributional text-based information and facilitate various downstream tasks. Since their manual construction is resource- and time-intensive, recent efforts have tried leveraging large pretrained language models (PLMs) to generate additional monolingual knowledge facts for KBs. However, such methods have not been attempted for building and enriching multilingual KBs. Besides wider application, such multilingual KBs can provide richer combined knowledge than monolingual (e.g., English) KBs. Knowledge expressed in different languages may be complementary and unequally distributed: this implies that the knowledge available in high-resource languages can be transferred to low-resource ones. To achieve this, it is crucial to represent multilingual knowledge in a shared/unified space. To this end, we propose a unified representation model, Prix-LM, for multilingual KB construction and completion. We leverage two types of knowledge, monolingual triples and cross-lingual links, extracted from existing multilingual KBs, and tune a multilingual language encoder XLM-R via a causal language modeling objective. Prix-LM integrates useful multilingual and KB-based factual knowledge into a single model. Experiments on standard entity-related tasks, such as link prediction in multiple languages, cross-lingual entity linking and bilingual lexicon induction, demonstrate its effectiveness, with gains reported over strong task-specialised baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 199d0527-dace-4a08-836b-5b7c263497edCited by top-tier papers2
- GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsDa Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li et al.EMNLP 2022 · 27 citations
- Enhancing Multilingual Language Model with Massive Multilingual Knowledge TriplesLinlin Liu, Xin Li, Ruidan He, Lidong Bing et al.EMNLP 2022 · 15 citations
Builds on8
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Knowledge Graph Alignment Network with Gated Multi-Hop Neighborhood AggregationZequn Sun, Chengming Wang, Wei Hu, Muhao Chen et al.AAAI 2020 · 379 citations
- Structure-Augmented Text Representation Learning for Efficient Knowledge Graph CompletionBo Wang, Tao Shen, Guodong Long, Tianyi Zhou et al.WWW 2021 · 322 citations
- A Benchmarking Study of Embedding-based Entity Alignment for Knowledge GraphsZequn Sun, Qingheng Zhang, Wei Hu, Chengming Wang et al.VLDB 2020 · 297 citations
- X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language ModelsZhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding et al.EMNLP 2020 · 81 citations
Related papers
- XLM-K: Improving Cross-Lingual Language Model Pre-training with Multilingual KnowledgeXiaoze Jiang, Yaobo Liang, Weizhu Chen, Nan DuanAAAI 2022 · 31 citations
- Multilingual Pre-training with Universal Dependency LearningKailai Sun, Zuchao Li, Hai ZhaoNeurIPS 2021 · 11 citations
- Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible MultilingualityMengyu Bu, Yang FengACL 2026 · 2 citations
- Entity Linking in 100 LanguagesJan A. Botha, Zifei Shan, Daniel GillickEMNLP 2020
- Massively Multilingual Lexical Specialization of Multilingual TransformersTommaso Green, Simone Paolo Ponzetto, Goran GlavasACL 2023
