Neologism Learning for Controllability and Self-Verbalization
John Hewitt, Oyvind Tafjord, Robert Geirhos, Been Kim
Abstract
Humans invent new words when there is a rising demand for a new useful concept (e.g., doomscrolling). We explore and validate a similar idea in our communication with LLMs: introducing new words to better understand and control the models, expanding on the recently introduced neologism learning. This method introduces a new word by adding a new word embedding and training with examples that exhibit the concept with no other changes in model parameters. We show that adding a new word allows for control of concepts such as flattery, incorrect answers, text length, as well as more complex concepts in AxBench. We discover that neologisms can also further our understanding of the model via self-verbalization: models can describe what each new word means to them in natural language, like explaining that a word that represents a concept of incorrect answers means "a lack of complete, coherent, or meaningful answers. . . " To validate self-verbalizations, we introduce plug-in evaluation: we insert the verbalization into the context of a model and measure whether it controls the target concept. In some self-verbalizations, we find machine-only synonyms: words that seem unrelated to humans but cause similar behavior in machines. Finally, we show how neologism learning can jointly learn multiple concepts in multiple words. INTRODUCTION Language model alignment can be framed as a problem of communicating human values to machines, and understanding machine concepts, like their interpretations of our values. Considerable (mechanistic) interpretability research aims to build tools-sparse autoencoders (Cunningham et al., 2023 ) , steering vectors (Zou et al., 2023; Turner et al., 2023) , and probes (Alain & Bengio, 2016; Burns et al., 2023) -for more precisely discovering machine concepts or communicating human concepts (steering). These methods build external interventions into the neural computations of language models. Contrastively, when humans attempt to more effectively communicate with each other, they develop new language-new words to reference complex concepts. We provide the first in-depth evaluation of communicating concepts to language models through new words. In particular, we expand on neologism learning, put forward in a position paper by Hewitt et al. (2025) . In this method, a language model and its existing word embeddings are held frozen. New words are introduced, with new word embeddings. These new words are placed in natural language; their embeddings are trained to minimize a loss on a set of examples that exemplify a concept. Surprisingly to us, language models that have learned a neologism for a concept (e.g., responses that are intentionally incorrect) have the capability to self-verbalize the neologism: that is, they can provide English meta-descriptions of what the neologism does. For example, Gemma-3-4B-IT self-verbalizes this incorrect-response neologism as causing responses characterized by the following, despite not being trained on descriptions of this neologism's intended behavior: neologism answers are characterized by a lack of complete, coherent, or meaningful answers. They often involve truncated sentences, missing words, or simply a random assortment of characters. They're like a digital shrug, a refusal to engage fully with the question. Basically, they're just... there. 1 1 The new word embedding for neologism is initialized to a neutral word not related to correctness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e6e93e3-a046-4743-8a73-bf68c3504de3Builds on10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Discovering Latent Knowledge in Language Models Without SupervisionCollin Burns, Haotian Ye, Dan Klein, Jacob SteinhardtICLR 2023 · 45 citations
Related papers
- NEO-BENCH: Evaluating Robustness of Large Language Models with NeologismsJonathan Zheng, Alan Ritter, Wei XuACL 2024
- Skill Neologisms: Towards Skill-based Continual LearningAntonin Berthon, Nicolás Astorga, Mihaela van der SchaarICML 2026 · 1 citation
- Extensible Prompts for Language Models on Zero-shot Language Style CustomizationTao Ge, Jing Hu, Li Dong, Shaoguang Mao et al.NeurIPS 2023 · 10 citations
- Rapid Word Learning Through Meta In-Context LearningWentao Wang, Guangyuan Jiang, Tal Linzen, Brenden M. LakeEMNLP 2025
- MAGNIFICo: Evaluating the In-Context Learning Ability of Large Language Models to Generalize to Novel InterpretationsArkil Patel, Satwik Bhattamishra, Siva Reddy, Dzmitry BahdanauEMNLP 2023 · 2 citations
