Fast WordPiece Tokenization
Xinying Song, Alex Salcianu, Yang Song, Dave Dopson, Denny Zhou
Abstract
Tokenization is a fundamental preprocessing step for almost all NLP tasks. In this paper, we propose efficient algorithms for the Word-Piece tokenization used in BERT, from singleword tokenization to general text (e.g., sentence) tokenization. When tokenizing a single word, WordPiece uses a longest-matchfirst strategy, known as maximum matching. The best known algorithms so far are ( 2 ) (where is the input length) or ( ) (where is the maximum vocabulary token length). We propose a novel algorithm whose tokenization complexity is strictly ( ). Our method is inspired by the Aho-Corasick algorithm. We introduce additional linkages on top of the trie built from the vocabulary, allowing smart transitions when the trie matching cannot continue. For general text, we further propose an algorithm that combines pre-tokenization (splitting the text into words) and our linear-time Word-Piece method into a single pass. Experimental results show that our method is 8.2x faster than HuggingFace Tokenizers and 5.1x faster than TensorFlow Text on average for general text tokenization. * Research conducted while working at Google.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0088e11d-08df-47c8-9334-3d114e259795Cited by top-tier papers19
- MAGNET: Improving the Multilingual Fairness of Language Models with Adaptive Gradient-Based TokenizationOrevaoghene Ahia, Sachin Kumar, Hila Gonen, Valentin Hofmann et al.NeurIPS 2024 · 37 citations
- Do All Languages Cost the Same? Tokenization in the Era of Commercial Language ModelsOrevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai et al.EMNLP 2023 · 24 citations
- StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-trainingYuechen Yu, Yulin Li, Chengquan Zhang, Xiaoqiang Zhang et al.ICLR 2023 · 18 citations
- Is Your LLM Overcharging You? Tokenization, Transparency, and IncentivesAnder Artola Velasco, Stratis Tsirtsis, Nastaran Okati, Manuel Gomez-RodriguezICML 2026 · 16 citations
- ScanDL: A Diffusion Model for Generating Synthetic Scanpaths on TextsLena S. Bolliger, David R. Reich, Patrick Haller, Deborah N. Jakobi et al.EMNLP 2023 · 6 citations
Builds on1
Related papers
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine et al.EMNLP 2024 · 16 citations
- A Partition Cover Approach to TokenizationJia Peng Lim, Shawn Tan, Davin Choo, Hady W. LauwNeurIPS 2025 · 6 citations
- Incremental BPE TokenizationShenghu Jiang, Ruihao GongICML 2026 · 12 citations
- ImagePiece: Content-aware Re-tokenization for Efficient Image RecognitionSeungdong Yoa, Seungjun Lee, Hye-Seung Cho, Bumsoo Kim et al.AAAI 2025 · 1 citation
- Pyramid-BERT: Reducing Complexity via Successive Core-set based Token SelectionXin Huang, Ashish Khetan, Rene Bidart, Zohar S. KarninACL 2022
