Evaluating Lexical Proficiency in Neural Language Models
Cristiano Ciaccio, Alessio Miaschi, Felice Dell'Orletta
摘要
We present a novel evaluation framework designed to assess the lexical proficiency and linguistic creativity of Transformer-based Language Models (LMs). We validate the framework by analyzing the performance of a set of LMs of different sizes, in both mono-and multilingual configuration, across tasks involving the generation, definition, and contextual usage of lexicalized words, neologisms, and nonce words. To support these evaluations, we developed a novel dataset of lexical entries for the Italian language, including curated definitions and usage examples sourced from various online platforms. The results highlight the robustness and effectiveness of our framework in evaluating multiple dimensions of LMs' linguistic understanding and offer an insight, through the assessment of their linguistic creativity, on the lexical generalization abilities of LMs 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Multi-Channel Reverse Dictionary ModelLei Zheng, Fanchao Qi, Zhiyuan Liu, Yasheng Wang 等AAAI 2020 · 被引用 46 次
- Generationary or "How We Went beyond Word Sense Inventories and Learned to Gloss"Michele Bevilacqua, Marco Maru, Roberto NavigliEMNLP 2020 · 被引用 41 次
- Character-Aware Models Improve Visual Text RenderingRosanne Liu, Dan Garrette, Chitwan Saharia, William Chan 等ACL 2023 · 被引用 25 次
- NEO-BENCH: Evaluating Robustness of Large Language Models with NeologismsJonathan Zheng, Alan Ritter, Wei XuACL 2024
相关 Paper
- Automated Creativity Evaluation of Language Models Across Open-Ended TasksTan Min Sen, Zachary Choy Kit Chun, Syed Ali Redha Alsagoff, Nadya Yuki Wangsajaya 等ACL 2026
- Evaluating Large Language Models via Linguistic ProfilingAlessio Miaschi, Felice Dell'Orletta, Giulia VenturiEMNLP 2024 · 被引用 1 次
- Death of the Novel(ty): Beyond N-Gram Novelty as a Metric for Textual CreativityArkadiy Saakyan, Najoung Kim, Smaranda Muresan, Tuhin ChakrabartyICLR 2026 · 被引用 6 次
- Exploring Precision and Recall to assess the quality and diversity of LLMsFlorian Le Bronnec, Alexandre Verine, Benjamin Négrevergne, Yann Chevaleyre 等ACL 2024 · 被引用 11 次
- Leveraging Large Language Models for NLG Evaluation: Advances and ChallengesZhen Li, Xiaohan Xu, Tao Shen, Can Xu 等EMNLP 2024 · 被引用 17 次
