StochasTok: Improving Fine-Grained Subword Understanding in LLMs
Anya Sims, Thomas Foster, T. Duy Nguyen-Hien, Klara Kaleb, Joseph Lee, Jakob Nicolaus Foerster, Yee Whye Teh, Cong Lu
Abstract
Subword-level understanding is integral to numerous tasks, including understanding multi-digit numbers, spelling mistakes, abbreviations, rhyming, and wordplay. Despite this, current large language models (LLMs) still struggle disproportionally with simple subword-level tasks like How many 'r's in 'strawberry'?. A key factor behind these failures is tokenization which obscures the fine-grained structure of words. Current alternatives, such as character-level and dropout tokenization methods, significantly increase computational costs and provide inconsistent improvements. In this paper we revisit tokenization and introduce STOCHASTOK, a simple, efficient stochastic tokenization scheme that randomly splits tokens during training, allowing LLMs to 'see' their internal structure. Our experiments show that pretraining with STOCHASTOK substantially improves LLMs' downstream performance across multiple subword-level language games, including character counting, substring identification, and math tasks. Furthermore, STOCHASTOK's simplicity allows seamless integration at any stage of the training pipeline; and we demonstrate that post-training with STOCHASTOK can instill improved subword understanding into existing pretrained models, thus avoiding costly pretraining from scratch. These dramatic improvements achieved with a minimal change suggest STOCHASTOK holds exciting potential when applied to larger, more capable models. Code open-sourced at: github.com/anyasims/stochastok.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e4126ec-22c6-4ba7-9377-9af8822a51eaCited by top-tier papers2
- Broken Tokens? Your Language Model can Secretly Handle Non-Canonical TokenizationsBrian Siyuan Zheng, Alisa Liu, Orevaoghene Ahia, Jonathan Hayase et al.NeurIPS 2025 · 19 citations
- SubTokenTest: A Practical Benchmark for Real-World Sub-token UnderstandingShuyang Hou, Yi Hu, Muhan ZhangACL 2026
Builds on7
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- MEGABYTE: Predicting Million-byte Sequences with Multiscale TransformersLili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan et al.NeurIPS 2023 · 197 citations
- Teaching Arithmetic to Small TransformersNayoung Lee, Kartik Sreenivasan, Jason D. Lee, Kangwook Lee et al.ICLR 2024 · 128 citations
- Byte Latent Transformer: Patches Scale Better Than TokensArtidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez, John Nguyen et al.ACL 2025 · 116 citations
Related papers
- CodeBPE: Investigating Subtokenization Options for Large Language Model Pretraining on Source CodeNadezhda Chirkova, Sergey TroshinICLR 2023 · 2 citations
- From Tokens to Words: On the Inner Lexicon of LLMsGuy Kaplan, Matanel Oren, Yuval Reif, Roy SchwartzICLR 2025
- Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language ModelsPit Neitemeier, Björn Deiseroth, Constantin Eichenberg, Lukas BallesICLR 2025
- CharBench: Evaluating the Role of Tokenization in Character-Level TasksOmri Uzan, Yuval PinterAAAI 2026 · 3 citations
- How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve ThemDisen Liao, Freda ShiACL 2026 · 1 citation
