Towards Building a Multilingual Sememe Knowledge Base: Predicting Sememes for BabelNet Synsets
Fanchao Qi, Liang Chang, Maosong Sun, Sicong Ouyang, Zhiyuan Liu
Abstract
A sememe is defined as the minimum semantic unit of human languages. Sememe knowledge bases (KBs), which contain words annotated with sememes, have been successfully applied to many NLP tasks. However, existing sememe KBs are built on only a few languages, which hinders their widespread utilization. To address the issue, we propose to build a unified sememe KB for multiple languages based on BabelNet, a multilingual encyclopedic dictionary. We first build a dataset serving as the seed of the multilingual sememe KB. It manually annotates sememes for over 15 thousand synsets (the entries of BabelNet). Then, we present a novel task of automatic sememe prediction for synsets, aiming to expand the seed dataset into a usable KB. We also propose two simple and effective models, which exploit different information of synsets. Finally, we conduct quantitative and qualitative analyses to explore important factors and difficulties in the task. All the source code and data of this work can be obtained on https://github.com/thunlp/BabelNet-Sememe-Prediction .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- How Sememic Components Can Benefit Link Prediction for Lexico-Semantic Knowledge Graphs?Hansi Wang, Yue Wang, Qiliang Liang, Yang LiuEMNLP 2025
- Enhancing Lexical Relation Mining with Structured Sememe KnowledgeHansi Wang, Qiliang Liang, Yue Wang, Yang LiuACL 2026
Related papers
- Massively Multilingual Lexical Specialization of Multilingual TransformersTommaso Green, Simone Paolo Ponzetto, Goran GlavasACL 2023
- Fully-Semantic Parsing and Generation: the BabelNet Meaning RepresentationAbelardo Carlos Martinez Lorenzo, Marco Maru, Roberto NavigliACL 2022 · 23 citations
- SensEmBERT: Context-Enhanced Sense Embeddings for Multilingual Word Sense DisambiguationBianca Scarlini, Tommaso Pasini, Roberto NavigliAAAI 2020 · 121 citations
- NewsEmbed: Modeling News through Pre-trained Document RepresentationsJialu Liu, Tianqi Liu, Cong YuKDD 2021 · 15 citations
- Improving Word Sense Disambiguation with TranslationsYixing Luan, Bradley Hauer, Lili Mou, Grzegorz KondrakEMNLP 2020 · 19 citations
