RoDEval: A Robust Word Sense Disambiguation Evaluation Framework for Large Language Models
Luyang Zhang, Shuaimin Li, Yishuo Li, Kunpeng Kang, Kaiyuan Zhang, Cong Wang, Wenpeng Lu
Abstract
Accurately evaluating the word sense disambiguation (WSD) capabilities of large language models (LLMs) remains challenging, as existing studies primarily rely on single-task evaluations and classification-based metrics that overlook the fundamental differences between generative LLMs and traditional classification models. To bridge this gap, we propose RoDEval, the first comprehensive evaluation framework specifically tailored for assessing LLMbased WSD methods. RoDEval introduces four novel metrics: Disambiguation Scope, Disambiguation Robustness, Disambiguation Reliability, and Definition Generation Quality Score, enabling a multifaceted evaluation of LLMs' WSD capabilities. Experimental results using RoDEval across five mainstream LLMs uncover significant limitations in their WSD performance. Specifically, incorrect definition selections in multiple-choice WSD tasks stem not from simple neglect or forget of correct options, but rather from incomplete acquisition of the all senses for polysemous words. Instead, disambiguation reliability is often compromised by the models' persistent overconfidence. In addition, inherent biases continue to affect performance, and scaling up model parameters alone fails to meaningfully enhance their ability to generate accurate sense definitions. These findings provide actionable insights for enhancing LLMs' WSD capabilities. The source code and evaluation scripts are open-sourced at https: //github.com/DayDream405/RoDEval .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- WSDPO: A Generative Word Sense Disambiguation Framework with Chain-of-Thought and Preference OptimizationKunpeng Kang, Shuaimin Li, Kaiyuan Zhang, Luyang Zhang et al.ACL 2026
- SOAPTriage: SOAP-Guided Multi-View Clinical Text Modeling Framework for Automated ESI PredictionEnming Wang, Jianlei Wang, Xueping Peng, Hongjiao Guan et al.ACL 2026
Builds on8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou et al.ICLR 2024 · 424 citations
- Extreme Compression of Large Language Models via Additive QuantizationVage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar et al.ICML 2024 · 187 citations
- Time Travel in LLMs: Tracing Data Contamination in Large Language ModelsShahriar Golchin, Mihai SurdeanuICLR 2024 · 165 citations
- Breaking Through the 80% Glass Ceiling: Raising the State of the Art in Word Sense Disambiguation by Incorporating Knowledge Graph InformationMichele Bevilacqua, Roberto NavigliACL 2020 · 145 citations
Related papers
- Do Large Language Models Understand Word Senses?Domenico Meconi, Simone Stirpe, Federico Martelli, Leonardo Lavalle et al.EMNLP 2025 · 7 citations
- Towards General-Domain Word Sense Disambiguation: Distilling Large Language Model into Compact DisambiguatorLiqiang Ming, Sheng-hua Zhong, Yuncong LiEMNLP 2025
- DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain TranslationZhibo Man, Yuanmeng Chen, Yujie Zhang, Jinan XuEMNLP 2025
- SenseRel: A Sense-Level Benchmark for Denotational and Connotational Meaning RelationsPierluigi Cassotti, Naomi Baes, Stefano De Pascale, Jáder Martins Camboim de Sá et al.ACL 2026
- F-Eval: Asssessing Fundamental Abilities with Refined Evaluation MethodsYu Sun, Keyuchen Keyuchen, Shujie Wang, Peiji Li et al.ACL 2024
