Large Scale Substitution-based Word Sense Induction
Matan Eyal, Shoval Sadde, Hillel Taub-Tabib, Yoav Goldberg
摘要
We present a word-sense induction method based on pre-trained masked language models (MLMs), which can cheaply scale to large vocabularies and large corpora. The result is a corpus which is sense-tagged according to a corpus-derived sense inventory and where each sense is associated with indicative words. Evaluation on English Wikipedia that was sense-tagged using our method shows that both the induced senses, and the per-instance sense assignment, are of high quality even compared to WSD methods, such as Babelfy. Furthermore, by training a static word embeddings algorithm on the sense-tagged corpus, we obtain high-quality static senseful embeddings. These outperform existing senseful embeddings methods on the WiC dataset and on a new outlier detection dataset we developed. The data driven nature of the algorithm allows to induce corpora-specific senses, which may not appear in standard sense inventories, as we demonstrate using a case study on the scientific domain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMsBenno Krojer, Perampalli Shravan Nayak, Oscar Mañas, Vaibhav Adlakha 等ICML 2026 · 被引用 6 次
- A Survey of Inductive Reasoning for Large Language ModelsKedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang 等ACL 2026 · 被引用 5 次
- What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classificationAndrew Halterman, Katherine A. KeithACL 2026 · 被引用 3 次
- To Word Senses and Beyond: Inducing Concepts with Contextualized Language ModelsBastien Liétard, Pascal Denis, Mikaela KellerEMNLP 2024
它引用的顶会 Paper1
相关 Paper
- Towards General-Domain Word Sense Disambiguation: Distilling Large Language Model into Compact DisambiguatorLiqiang Ming, Sheng-hua Zhong, Yuncong LiEMNLP 2025
- CluBERT: A Cluster-Based Approach for Learning Sense Distributions in Multiple LanguagesTommaso Pasini, Federico Scozzafava, Bianca ScarliniACL 2020 · 被引用 23 次
- SenseBERT: Driving Some Sense into BERTYoav Levine, Barak Lenz, Or Dagan, Ori Ram 等ACL 2020 · 被引用 27 次
- Towards Semantics-Enhanced Pre-Training: Can Lexicon Definitions Help Learning Sentence Meanings?Xuancheng Ren, Xu Sun, Houfeng Wang, Qun LiuAAAI 2021 · 被引用 5 次
- FOOL ME IF YOU CAN! An Adversarial Dataset to Investigate the Robustness of LMs in Word Sense DisambiguationMohamad Ballout, Anne Dedert, Nohayr Abdelmoneim, Ulf Krumnack 等EMNLP 2024 · 被引用 2 次
