Nibbling at the Hard Core of Word Sense Disambiguation
Marco Maru, Simone Conia, Michele Bevilacqua, Roberto Navigli
摘要
With state-of-the-art systems having finally attained estimated human performance, Word Sense Disambiguation (WSD) has now joined the array of Natural Language Processing tasks that have seemingly been solved, thanks to the vast amounts of knowledge encoded into Transformer-based pre-trained language models. And yet, if we look below the surface of raw figures, it is easy to realize that current approaches still make trivial mistakes that a human would never make. In this work, we provide evidence showing why the F1 score metric should not simply be taken at face value and present an exhaustive analysis of the errors that seven of the most representative state-of-the-art systems for English all-words WSD make on traditional evaluation benchmarks. In addition, we produce and release a collection of test sets featuring (a) an amended version of the standard evaluation benchmark that fixes its lexical and semantic inaccuracies, (b) 42D, a challenge set devised to assess the resilience of systems with respect to least frequent word senses and senses not seen at training time, and (c) hardEN, a challenge set made up solely of instances which none of the investigated state-of-the-art systems can solve. We make all of the test sets and model predictions available to the research community at https://github.com/ SapienzaNLP/wsd-hard-benchmark .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Do Large Language Models Understand Word Senses?Domenico Meconi, Simone Stirpe, Federico Martelli, Leonardo Lavalle 等EMNLP 2025 · 被引用 7 次
- Increasing Coverage and Precision of Textual Information in Multilingual Knowledge GraphsSimone Conia, Min Li, Daniel Lee, Umar Farooq Minhas 等EMNLP 2023 · 被引用 3 次
- Silent Signals, Loud Impact: LLMs for Word-Sense Disambiguation of Coded Dog WhistlesJulia Kruk, Michela Marchini, Rijul Magu, Caleb Ziems 等ACL 2024 · 被引用 2 次
- WSDPO: A Generative Word Sense Disambiguation Framework with Chain-of-Thought and Preference OptimizationKunpeng Kang, Shuaimin Li, Kaiyuan Zhang, Luyang Zhang 等ACL 2026
- Towards General-Domain Word Sense Disambiguation: Distilling Large Language Model into Compact DisambiguatorLiqiang Ming, Sheng-hua Zhong, Yuncong LiEMNLP 2025
它引用的顶会 Paper12
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Breaking Through the 80% Glass Ceiling: Raising the State of the Art in Word Sense Disambiguation by Incorporating Knowledge Graph InformationMichele Bevilacqua, Roberto NavigliACL 2020 · 被引用 145 次
- With More Contexts Comes Better Performance: Contextualized Sense Embeddings for All-Round Word Sense DisambiguationBianca Scarlini, Tommaso Pasini, Roberto NavigliEMNLP 2020 · 被引用 95 次
- With Little Power Comes Great ResponsibilityDallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia 等EMNLP 2020 · 被引用 76 次
- XL-WSD: An Extra-Large and Cross-Lingual Evaluation Framework for Word Sense DisambiguationTommaso Pasini, Alessandro Raganato, Roberto NavigliAAAI 2021 · 被引用 76 次
相关 Paper
- ConSeC: Word Sense Disambiguation as Continuous Sense ComprehensionEdoardo Barba, Luigi Procopio, Roberto NavigliEMNLP 2021 · 被引用 60 次
- FOOL ME IF YOU CAN! An Adversarial Dataset to Investigate the Robustness of LMs in Word Sense DisambiguationMohamad Ballout, Anne Dedert, Nohayr Abdelmoneim, Ulf Krumnack 等EMNLP 2024 · 被引用 2 次
- RoDEval: A Robust Word Sense Disambiguation Evaluation Framework for Large Language ModelsLuyang Zhang, Shuaimin Li, Yishuo Li, Kunpeng Kang 等EMNLP 2025
- How Much Do Encoder Models Know About Word Senses?Simone Teglia, Simone Tedeschi, Roberto NavigliACL 2025 · 被引用 1 次
- Back Deduction Based Testing for Word Sense Disambiguation Ability of Machine Translation SystemsJun Wang, Yanhui Li, Xiang Huang, Lin Chen 等ISSTA 2023 · 被引用 4 次
