NEO-BENCH: Evaluating Robustness of Large Language Models with Neologisms
Jonathan Zheng, Alan Ritter, Wei Xu
摘要
The performance of Large Language Models (LLMs) degrades from the temporal drift between data used for model training and newer text seen during inference. One understudied avenue of language change causing data drift is the emergence of neologisms -new word forms -over time. We create a diverse resource of recent English neologisms by using several popular collection methods. We analyze temporal drift using neologisms by comparing sentences containing new words with near-identical sentences that replace neologisms with existing substitute words. Model performance is nearly halved in machine translation when a single neologism is introduced in a sentence. Motivated by these results, we construct a benchmark to evaluate LLMs' ability to generalize to neologisms with various natural language understanding tasks and model perplexity. Models with later knowledge cutoff dates yield lower perplexities and perform better in downstream tasks. LLMs are also affected differently based on the linguistic origins of words, indicating that neologisms are complex for static LLMs to address. We will release our benchmark and code for reproducing our experiments. BART T5 GPT-J GPT-3.5 GPT-4 LLaMA-1 LLaMA-2 Pig Butchering Maskne Barbiecore Figure 1: NEO-BENCH collects neologisms from 2020-2023 for LLM evaluation. "Pig Butchering" originated as a Mandarin expression (杀猪盘). wal and Nenkova, 2022; Liu and Ritter, 2023). 042 However, as far as we are aware there has not been 043 prior work that analyzes the robustness of LLMs on 044 handling neologisms. We show that adding a neol-045 ogism to text decreases machine translation quality 046 by an average of 44% in a human evaluation ( §2),
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- CodeAgent: Autonomous Communicative Agents for Code ReviewXunzhu Tang, Kisub Kim, Yewei Song, Cedric Lothritz 等EMNLP 2024 · 被引用 8 次
- NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement LearningZhongtao Miao, Kaiyan Zhao, Masaaki Nagata, Yoshimasa TsuruokaACL 2026 · 被引用 2 次
- DeReA: Improving Idiom Translation with Detect-Retrieve-Arbitrate ReasoningRongqing Jiang, Xuebo Liu, Shengxin Liu, Yutong Wang 等ACL 2026
- Evaluating Lexical Proficiency in Neural Language ModelsCristiano Ciaccio, Alessio Miaschi, Felice Dell'OrlettaACL 2025
- SDBench: A Survey-based Domain-specific LLM Benchmarking and Optimization FrameworkCheng Guo, Hu Kai, Shuxian Liang, Yiyang Jiang 等ACL 2025
它引用的顶会 Paper2
相关 Paper
- Neologism Learning for Controllability and Self-VerbalizationJohn Hewitt, Oyvind Tafjord, Robert Geirhos, Been KimICLR 2026 · 被引用 6 次
- Improving Temporal Generalization of Pre-trained Language Models with Lexical Semantic ChangeZhaochen Su, Zecheng Tang, Xinyan Guan, Lijun Wu 等EMNLP 2022 · 被引用 11 次
- Are LLMs Prescient? A Continuous Evaluation using Daily News as the OracleHui Dai, Ryan Teehan, Mengye RenICML 2025
- CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 LanguagesYilun Yang, Yekun ChaiEMNLP 2025 · 被引用 1 次
- New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMsShiyao Cui, Qinglin Zhang, Di Wang, Yida Lu 等ACL 2026
