NEO-BENCH: Evaluating Robustness of Large Language Models with Neologisms
Jonathan Zheng, Alan Ritter, Wei Xu
Abstract
The performance of Large Language Models (LLMs) degrades from the temporal drift between data used for model training and newer text seen during inference. One understudied avenue of language change causing data drift is the emergence of neologisms -new word forms -over time. We create a diverse resource of recent English neologisms by using several popular collection methods. We analyze temporal drift using neologisms by comparing sentences containing new words with near-identical sentences that replace neologisms with existing substitute words. Model performance is nearly halved in machine translation when a single neologism is introduced in a sentence. Motivated by these results, we construct a benchmark to evaluate LLMs' ability to generalize to neologisms with various natural language understanding tasks and model perplexity. Models with later knowledge cutoff dates yield lower perplexities and perform better in downstream tasks. LLMs are also affected differently based on the linguistic origins of words, indicating that neologisms are complex for static LLMs to address. We will release our benchmark and code for reproducing our experiments. BART T5 GPT-J GPT-3.5 GPT-4 LLaMA-1 LLaMA-2 Pig Butchering Maskne Barbiecore Figure 1: NEO-BENCH collects neologisms from 2020-2023 for LLM evaluation. "Pig Butchering" originated as a Mandarin expression (杀猪盘). wal and Nenkova, 2022; Liu and Ritter, 2023). 042 However, as far as we are aware there has not been 043 prior work that analyzes the robustness of LLMs on 044 handling neologisms. We show that adding a neol-045 ogism to text decreases machine translation quality 046 by an average of 44% in a human evaluation ( §2),
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e056b77-4a66-42ce-8527-6b3553db26d5Cited by top-tier papers7
- CodeAgent: Autonomous Communicative Agents for Code ReviewXunzhu Tang, Kisub Kim, Yewei Song, Cedric Lothritz et al.EMNLP 2024 · 8 citations
- NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement LearningZhongtao Miao, Kaiyan Zhao, Masaaki Nagata, Yoshimasa TsuruokaACL 2026 · 2 citations
- DeReA: Improving Idiom Translation with Detect-Retrieve-Arbitrate ReasoningRongqing Jiang, Xuebo Liu, Shengxin Liu, Yutong Wang et al.ACL 2026
- Evaluating Lexical Proficiency in Neural Language ModelsCristiano Ciaccio, Alessio Miaschi, Felice Dell'OrlettaACL 2025
- SDBench: A Survey-based Domain-specific LLM Benchmarking and Optimization FrameworkCheng Guo, Hu Kai, Shuxian Liang, Yiyang Jiang et al.ACL 2025
Builds on2
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language ModelsJoel Jang, Seonghyeon Ye, Changho Lee, Sohee Yang et al.EMNLP 2022 · 42 citations
Related papers
- Neologism Learning for Controllability and Self-VerbalizationJohn Hewitt, Oyvind Tafjord, Robert Geirhos, Been KimICLR 2026 · 6 citations
- Improving Temporal Generalization of Pre-trained Language Models with Lexical Semantic ChangeZhaochen Su, Zecheng Tang, Xinyan Guan, Lijun Wu et al.EMNLP 2022 · 11 citations
- Are LLMs Prescient? A Continuous Evaluation using Daily News as the OracleHui Dai, Ryan Teehan, Mengye RenICML 2025
- CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 LanguagesYilun Yang, Yekun ChaiEMNLP 2025 · 1 citation
- New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMsShiyao Cui, Qinglin Zhang, Di Wang, Yida Lu et al.ACL 2026
