RoDEval: A Robust Word Sense Disambiguation Evaluation Framework for Large Language Models
Luyang Zhang, Shuaimin Li, Yishuo Li, Kunpeng Kang, Kaiyuan Zhang, Cong Wang, Wenpeng Lu
摘要
Accurately evaluating the word sense disambiguation (WSD) capabilities of large language models (LLMs) remains challenging, as existing studies primarily rely on single-task evaluations and classification-based metrics that overlook the fundamental differences between generative LLMs and traditional classification models. To bridge this gap, we propose RoDEval, the first comprehensive evaluation framework specifically tailored for assessing LLMbased WSD methods. RoDEval introduces four novel metrics: Disambiguation Scope, Disambiguation Robustness, Disambiguation Reliability, and Definition Generation Quality Score, enabling a multifaceted evaluation of LLMs' WSD capabilities. Experimental results using RoDEval across five mainstream LLMs uncover significant limitations in their WSD performance. Specifically, incorrect definition selections in multiple-choice WSD tasks stem not from simple neglect or forget of correct options, but rather from incomplete acquisition of the all senses for polysemous words. Instead, disambiguation reliability is often compromised by the models' persistent overconfidence. In addition, inherent biases continue to affect performance, and scaling up model parameters alone fails to meaningfully enhance their ability to generate accurate sense definitions. These findings provide actionable insights for enhancing LLMs' WSD capabilities. The source code and evaluation scripts are open-sourced at https: //github.com/DayDream405/RoDEval .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- WSDPO: A Generative Word Sense Disambiguation Framework with Chain-of-Thought and Preference OptimizationKunpeng Kang, Shuaimin Li, Kaiyuan Zhang, Luyang Zhang 等ACL 2026
- SOAPTriage: SOAP-Guided Multi-View Clinical Text Modeling Framework for Automated ESI PredictionEnming Wang, Jianlei Wang, Xueping Peng, Hongjiao Guan 等ACL 2026
它引用的顶会 Paper8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou 等ICLR 2024 · 被引用 424 次
- Extreme Compression of Large Language Models via Additive QuantizationVage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar 等ICML 2024 · 被引用 187 次
- Time Travel in LLMs: Tracing Data Contamination in Large Language ModelsShahriar Golchin, Mihai SurdeanuICLR 2024 · 被引用 165 次
- Breaking Through the 80% Glass Ceiling: Raising the State of the Art in Word Sense Disambiguation by Incorporating Knowledge Graph InformationMichele Bevilacqua, Roberto NavigliACL 2020 · 被引用 145 次
相关 Paper
- Do Large Language Models Understand Word Senses?Domenico Meconi, Simone Stirpe, Federico Martelli, Leonardo Lavalle 等EMNLP 2025 · 被引用 7 次
- Towards General-Domain Word Sense Disambiguation: Distilling Large Language Model into Compact DisambiguatorLiqiang Ming, Sheng-hua Zhong, Yuncong LiEMNLP 2025
- DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain TranslationZhibo Man, Yuanmeng Chen, Yujie Zhang, Jinan XuEMNLP 2025
- SenseRel: A Sense-Level Benchmark for Denotational and Connotational Meaning RelationsPierluigi Cassotti, Naomi Baes, Stefano De Pascale, Jáder Martins Camboim de Sá 等ACL 2026
- F-Eval: Asssessing Fundamental Abilities with Refined Evaluation MethodsYu Sun, Keyuchen Keyuchen, Shujie Wang, Peiji Li 等ACL 2024
