Lune

ICLR2026顶会

StochasTok: Improving Fine-Grained Subword Understanding in LLMs

Anya Sims, Thomas Foster, T. Duy Nguyen-Hien, Klara Kaleb, Joseph Lee, Jakob Nicolaus Foerster, Yee Whye Teh, Cong Lu

2026年份
8被引次数
2顶会引用

摘要

Subword-level understanding is integral to numerous tasks, including understanding multi-digit numbers, spelling mistakes, abbreviations, rhyming, and wordplay. Despite this, current large language models (LLMs) still struggle disproportionally with simple subword-level tasks like How many 'r's in 'strawberry'?. A key factor behind these failures is tokenization which obscures the fine-grained structure of words. Current alternatives, such as character-level and dropout tokenization methods, significantly increase computational costs and provide inconsistent improvements. In this paper we revisit tokenization and introduce STOCHASTOK, a simple, efficient stochastic tokenization scheme that randomly splits tokens during training, allowing LLMs to 'see' their internal structure. Our experiments show that pretraining with STOCHASTOK substantially improves LLMs' downstream performance across multiple subword-level language games, including character counting, substring identification, and math tasks. Furthermore, STOCHASTOK's simplicity allows seamless integration at any stage of the training pipeline; and we demonstrate that post-training with STOCHASTOK can instill improved subword understanding into existing pretrained models, thus avoiding costly pretraining from scratch. These dramatic improvements achieved with a minimal change suggest STOCHASTOK holds exciting potential when applied to larger, more capable models. Code open-sourced at: github.com/anyasims/stochastok.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper7

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖