CluBERT: A Cluster-Based Approach for Learning Sense Distributions in Multiple Languages
Tommaso Pasini, Federico Scozzafava, Bianca Scarlini
Abstract
Knowing the Most Frequent Sense (MFS) of a word has been proved to help Word Sense Disambiguation (WSD) models significantly. However, the scarcity of sense-annotated data makes it difficult to induce a reliable and highcoverage distribution of the meanings in a language vocabulary. To address this issue, in this paper we present CluBERT, an automatic and multilingual approach for inducing the distributions of word senses from a corpus of raw sentences. Our experiments show that Clu-BERT learns distributions over English senses that are of higher quality than those extracted by alternative approaches. When used to induce the MFS of a lemma, CluBERT attains state-of-the-art results on the English Word Sense Disambiguation tasks and helps to improve the disambiguation performance of two off-the-shelf WSD models. Moreover, our distributions also prove to be effective in other languages, beating all their alternatives for computing the MFS on the multilingual WSD tasks. We release our sense distributions in five different languages at https://github. com/SapienzaNLP/clubert .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 012bf1ca-08fd-4be4-8617-0ba268fcf523Cited by top-tier papers4
- With More Contexts Comes Better Performance: Contextualized Sense Embeddings for All-Round Word Sense DisambiguationBianca Scarlini, Tommaso Pasini, Roberto NavigliEMNLP 2020 · 95 citations
- XL-WSD: An Extra-Large and Cross-Lingual Evaluation Framework for Word Sense DisambiguationTommaso Pasini, Alessandro Raganato, Roberto NavigliAAAI 2021 · 76 citations
- Improving Word Sense Disambiguation with TranslationsYixing Luan, Bradley Hauer, Lili Mou, Grzegorz KondrakEMNLP 2020 · 19 citations
- Pre-training and Fine-tuning Neural Topic Model: A Simple yet Effective Approach to Incorporating External KnowledgeLinhai Zhang, Xuemeng Hu, Boyu Wang, Deyu Zhou et al.ACL 2022 · 14 citations
Builds on1
Related papers
- How Much Do Encoder Models Know About Word Senses?Simone Teglia, Simone Tedeschi, Roberto NavigliACL 2025 · 1 citation
- SenseBERT: Driving Some Sense into BERTYoav Levine, Barak Lenz, Or Dagan, Ori Ram et al.ACL 2020 · 27 citations
- Rare and Zero-shot Word Sense Disambiguation using Z-ReweightingYing Su, Hongming Zhang, Yangqiu Song, Tong ZhangACL 2022 · 12 citations
- Towards General-Domain Word Sense Disambiguation: Distilling Large Language Model into Compact DisambiguatorLiqiang Ming, Sheng-hua Zhong, Yuncong LiEMNLP 2025
- Large Scale Substitution-based Word Sense InductionMatan Eyal, Shoval Sadde, Hillel Taub-Tabib, Yoav GoldbergACL 2022
