From Tokens to Words: On the Inner Lexicon of LLMs
Guy Kaplan, Matanel Oren, Yuval Reif, Roy Schwartz
摘要
Natural language is composed of words, but modern large language models (LLMs) process sub-words as input. A natural question raised by this discrepancy is whether LLMs encode words internally, and if so how. We present evidence that LLMs engage in an intrinsic detokenization process, where subword sequences are combined into coherent whole-word representations at their last token. Our experiments show that this process primarily takes place within the early and middle layers of the model. We further demonstrate its robustness to arbitrary splits (e.g., "cats" to "ca" and "ts"), typos, and importantly-to outof-vocabulary words: when feeding the last token internal representations of such words to the model as input, it can "understand" them as the complete word despite never seeing such representations as input during training. Our findings suggest that LLMs maintain a latent vocabulary beyond the tokenizer's scope. These insights provide a practical, finetuning-free application for expanding the vocabulary of pre-trained models. By enabling the addition of new vocabulary words, we reduce input length and inference iterations, which reduces both space and model latency, with little to no loss in model accuracy. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Broken Tokens? Your Language Model can Secretly Handle Non-Canonical TokenizationsBrian Siyuan Zheng, Alisa Liu, Orevaoghene Ahia, Jonathan Hayase 等NeurIPS 2025 · 被引用 19 次
- Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization FailureBoshi Wang, Huan SunICLR 2026 · 被引用 16 次
- The Strawberry Problem: Emergence of Character-level Understanding in Tokenized Language ModelsAdrian Cosma, Stefan Ruseti, Emilian Radoi, Mihai DascaluEMNLP 2025 · 被引用 13 次
- Measuring and Guiding MonosemanticityRuben Härle, Felix Friedrich, Manuel Brack, Björn Deiseroth 等NeurIPS 2025 · 被引用 12 次
- Elevating Visual Perception in Multimodal LLMs with Visual Embedding DistillationJitesh Jain, Zhengyuan Yang, Humphrey Shi, Jianfeng Gao 等NeurIPS 2025 · 被引用 9 次
它引用的顶会 Paper20
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Language Model Tokenizers Introduce Unfairness Between LanguagesAleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, Adel BibiNeurIPS 2023 · 被引用 301 次
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 被引用 233 次
- Function Vectors in Large Language ModelsEric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller 等ICLR 2024 · 被引用 229 次
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsAsma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon 等ICML 2024 · 被引用 197 次
相关 Paper
- From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language UnderstandingLi Sun, Florian Luisier, Kayhan Batmanghelich, Dinei A. F. Florêncio 等ACL 2023
- StochasTok: Improving Fine-Grained Subword Understanding in LLMsAnya Sims, Thomas Foster, T. Duy Nguyen-Hien, Klara Kaleb 等ICLR 2026 · 被引用 8 次
- Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language ModelsPit Neitemeier, Björn Deiseroth, Constantin Eichenberg, Lukas BallesICLR 2025
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM HallucinationsHadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart 等ICLR 2025
- Understanding Subword Compositionality of Large Language ModelsQiwei Peng, Yekun Chai, Anders SøgaardEMNLP 2025
