The Strawberry Problem: Emergence of Character-level Understanding in Tokenized Language Models
Adrian Cosma, Stefan Ruseti, Emilian Radoi, Mihai Dascalu
Abstract
Despite their remarkable progress across diverse domains, Large Language Models (LLMs) consistently fail at simple character-level tasks, such as counting letters in words, due to a fundamental limitation: tokenization. In this work, we frame this limitation as a problem of low mutual information and analyze it in terms of concept emergence. Using a suite of 19 synthetic tasks that isolate character-level reasoning in a controlled setting, we show that such capabilities emerge suddenly and only late in training. We find that percolation-based models of concept emergence explain these patterns, suggesting that learning character composition is not fundamentally different from learning commonsense knowledge. To address this bottleneck, we propose a lightweight architectural modification that significantly improves character-level reasoning while preserving the inductive advantages of subword models. Together, our results bridge low-level perceptual gaps in tokenized LMs and provide a principled framework for understanding and mitigating their structural blind spots. We make our code publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8435be18-7512-4e07-8db3-b6505e781648Cited by top-tier papers1
Ask how each one uses itBuilds on8
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 796 citations
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando et al.ICML 2023 · 474 citations
- Charformer: Fast Character Transformers via Gradient-based Subword TokenizationYi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta et al.ICLR 2022 · 198 citations
- Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept SpaceCore Francisco Park, Maya Okawa, Andrew Lee, Ekdeep Singh Lubana et al.NeurIPS 2024 · 39 citations
Related papers
- SubTokenTest: A Practical Benchmark for Real-World Sub-token UnderstandingShuyang Hou, Yi Hu, Muhan ZhangACL 2026
- CharBench: Evaluating the Role of Tokenization in Character-Level TasksOmri Uzan, Yuval PinterAAAI 2026 · 3 citations
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- Emergent Hierarchical Reasoning in LLMs through Reinforcement LearningHaozhe Wang, Qixin Xu, Che Liu, Junhong Wu et al.ICLR 2026 · 44 citations
- Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language ModelsPit Neitemeier, Björn Deiseroth, Constantin Eichenberg, Lukas BallesICLR 2025
