Nibbling at the Hard Core of Word Sense Disambiguation
Marco Maru, Simone Conia, Michele Bevilacqua, Roberto Navigli
Abstract
With state-of-the-art systems having finally attained estimated human performance, Word Sense Disambiguation (WSD) has now joined the array of Natural Language Processing tasks that have seemingly been solved, thanks to the vast amounts of knowledge encoded into Transformer-based pre-trained language models. And yet, if we look below the surface of raw figures, it is easy to realize that current approaches still make trivial mistakes that a human would never make. In this work, we provide evidence showing why the F1 score metric should not simply be taken at face value and present an exhaustive analysis of the errors that seven of the most representative state-of-the-art systems for English all-words WSD make on traditional evaluation benchmarks. In addition, we produce and release a collection of test sets featuring (a) an amended version of the standard evaluation benchmark that fixes its lexical and semantic inaccuracies, (b) 42D, a challenge set devised to assess the resilience of systems with respect to least frequent word senses and senses not seen at training time, and (c) hardEN, a challenge set made up solely of instances which none of the investigated state-of-the-art systems can solve. We make all of the test sets and model predictions available to the research community at https://github.com/ SapienzaNLP/wsd-hard-benchmark .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Do Large Language Models Understand Word Senses?Domenico Meconi, Simone Stirpe, Federico Martelli, Leonardo Lavalle et al.EMNLP 2025 · 7 citations
- Increasing Coverage and Precision of Textual Information in Multilingual Knowledge GraphsSimone Conia, Min Li, Daniel Lee, Umar Farooq Minhas et al.EMNLP 2023 · 3 citations
- Silent Signals, Loud Impact: LLMs for Word-Sense Disambiguation of Coded Dog WhistlesJulia Kruk, Michela Marchini, Rijul Magu, Caleb Ziems et al.ACL 2024 · 2 citations
- WSDPO: A Generative Word Sense Disambiguation Framework with Chain-of-Thought and Preference OptimizationKunpeng Kang, Shuaimin Li, Kaiyuan Zhang, Luyang Zhang et al.ACL 2026
- Towards General-Domain Word Sense Disambiguation: Distilling Large Language Model into Compact DisambiguatorLiqiang Ming, Sheng-hua Zhong, Yuncong LiEMNLP 2025
Builds on12
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Breaking Through the 80% Glass Ceiling: Raising the State of the Art in Word Sense Disambiguation by Incorporating Knowledge Graph InformationMichele Bevilacqua, Roberto NavigliACL 2020 · 145 citations
- With More Contexts Comes Better Performance: Contextualized Sense Embeddings for All-Round Word Sense DisambiguationBianca Scarlini, Tommaso Pasini, Roberto NavigliEMNLP 2020 · 95 citations
- With Little Power Comes Great ResponsibilityDallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia et al.EMNLP 2020 · 76 citations
- XL-WSD: An Extra-Large and Cross-Lingual Evaluation Framework for Word Sense DisambiguationTommaso Pasini, Alessandro Raganato, Roberto NavigliAAAI 2021 · 76 citations
Related papers
- ConSeC: Word Sense Disambiguation as Continuous Sense ComprehensionEdoardo Barba, Luigi Procopio, Roberto NavigliEMNLP 2021 · 60 citations
- FOOL ME IF YOU CAN! An Adversarial Dataset to Investigate the Robustness of LMs in Word Sense DisambiguationMohamad Ballout, Anne Dedert, Nohayr Abdelmoneim, Ulf Krumnack et al.EMNLP 2024 · 2 citations
- RoDEval: A Robust Word Sense Disambiguation Evaluation Framework for Large Language ModelsLuyang Zhang, Shuaimin Li, Yishuo Li, Kunpeng Kang et al.EMNLP 2025
- How Much Do Encoder Models Know About Word Senses?Simone Teglia, Simone Tedeschi, Roberto NavigliACL 2025 · 1 citation
- Back Deduction Based Testing for Word Sense Disambiguation Ability of Machine Translation SystemsJun Wang, Yanhui Li, Xiang Huang, Lin Chen et al.ISSTA 2023 · 4 citations
