Towards General-Domain Word Sense Disambiguation: Distilling Large Language Model into Compact Disambiguator
Liqiang Ming, Sheng-hua Zhong, Yuncong Li
Abstract
Word Sense Disambiguation (WSD) aims to determine the correct meaning of a word in context from a predefined inventory, and remains a fundamental challenge in natural language understanding. Existing methods rely heavily on manually annotated data, which limits coverage and generalization. In this work, we propose a scalable framework that leverages large language models (LLMs) as knowledge distillers to construct silver-standard WSD corpora. We explore generation-based distillation, where diverse examples are synthesized for dictionary senses, and annotation-based distillation, where LLMs assign sense labels to polysemous words within real-world corpus sentences. The resulting data is used to train tiny models. Extensive experiments show that models distilled from LLM-generated data outperform those trained on gold-standard corpora, especially on general-domain benchmarks. Our annotationbased model, after balancing sense distribution, achieves 50% F1 gain on the most challenging test set and the best distilled model can match or even exceed the performance of its LLM teacher, despite having over 1000 times fewer parameters. These results demonstrate the effectiveness of LLM-based distillation for building accurate, generalizable, and efficient WSD systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c04663be-440d-41c4-a38e-3418e5ee312dCited by top-tier papers1
Ask how each one uses itBuilds on10
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Generating Training Data with Language Models: Towards Zero-Shot Language UnderstandingYu Meng, Jiaxin Huang, Yu Zhang, Jiawei HanNeurIPS 2022 · 309 citations
- UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionWenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen et al.ICLR 2024 · 118 citations
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu et al.EMNLP 2022 · 96 citations
- ConSeC: Word Sense Disambiguation as Continuous Sense ComprehensionEdoardo Barba, Luigi Procopio, Roberto NavigliEMNLP 2021 · 60 citations
Related papers
- RoDEval: A Robust Word Sense Disambiguation Evaluation Framework for Large Language ModelsLuyang Zhang, Shuaimin Li, Yishuo Li, Kunpeng Kang et al.EMNLP 2025
- Do Large Language Models Understand Word Senses?Domenico Meconi, Simone Stirpe, Federico Martelli, Leonardo Lavalle et al.EMNLP 2025 · 7 citations
- XL-WSD: An Extra-Large and Cross-Lingual Evaluation Framework for Word Sense DisambiguationTommaso Pasini, Alessandro Raganato, Roberto NavigliAAAI 2021 · 76 citations
- MADAWSD: Multi-Agent Debate Framework for Adversarial Word Sense DisambiguationKaiyuan Zhang, Qian Liu, Luyang Zhang, Chaoqun Zheng et al.EMNLP 2025
- Large Scale Substitution-based Word Sense InductionMatan Eyal, Shoval Sadde, Hillel Taub-Tabib, Yoav GoldbergACL 2022
