Fast WordPiece Tokenization
Xinying Song, Alex Salcianu, Yang Song, Dave Dopson, Denny Zhou
摘要
Tokenization is a fundamental preprocessing step for almost all NLP tasks. In this paper, we propose efficient algorithms for the Word-Piece tokenization used in BERT, from singleword tokenization to general text (e.g., sentence) tokenization. When tokenizing a single word, WordPiece uses a longest-matchfirst strategy, known as maximum matching. The best known algorithms so far are ( 2 ) (where is the input length) or ( ) (where is the maximum vocabulary token length). We propose a novel algorithm whose tokenization complexity is strictly ( ). Our method is inspired by the Aho-Corasick algorithm. We introduce additional linkages on top of the trie built from the vocabulary, allowing smart transitions when the trie matching cannot continue. For general text, we further propose an algorithm that combines pre-tokenization (splitting the text into words) and our linear-time Word-Piece method into a single pass. Experimental results show that our method is 8.2x faster than HuggingFace Tokenizers and 5.1x faster than TensorFlow Text on average for general text tokenization. * Research conducted while working at Google.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- MAGNET: Improving the Multilingual Fairness of Language Models with Adaptive Gradient-Based TokenizationOrevaoghene Ahia, Sachin Kumar, Hila Gonen, Valentin Hofmann 等NeurIPS 2024 · 被引用 37 次
- Do All Languages Cost the Same? Tokenization in the Era of Commercial Language ModelsOrevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai 等EMNLP 2023 · 被引用 24 次
- StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-trainingYuechen Yu, Yulin Li, Chengquan Zhang, Xiaoqiang Zhang 等ICLR 2023 · 被引用 18 次
- Is Your LLM Overcharging You? Tokenization, Transparency, and IncentivesAnder Artola Velasco, Stratis Tsirtsis, Nastaran Okati, Manuel Gomez-RodriguezICML 2026 · 被引用 16 次
- ScanDL: A Diffusion Model for Generating Synthetic Scanpaths on TextsLena S. Bolliger, David R. Reich, Patrick Haller, Deborah N. Jakobi 等EMNLP 2023 · 被引用 6 次
它引用的顶会 Paper1
相关 Paper
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine 等EMNLP 2024 · 被引用 16 次
- A Partition Cover Approach to TokenizationJia Peng Lim, Shawn Tan, Davin Choo, Hady W. LauwNeurIPS 2025 · 被引用 6 次
- Incremental BPE TokenizationShenghu Jiang, Ruihao GongICML 2026 · 被引用 12 次
- ImagePiece: Content-aware Re-tokenization for Efficient Image RecognitionSeungdong Yoa, Seungjun Lee, Hye-Seung Cho, Bumsoo Kim 等AAAI 2025 · 被引用 1 次
- Pyramid-BERT: Reducing Complexity via Successive Core-set based Token SelectionXin Huang, Ashish Khetan, Rene Bidart, Zohar S. KarninACL 2022
