BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
Pavel Chizhov, Catherine Arnett, Elizaveta Korotkova, Ivan P. Yamshchikov
Abstract
Language models can greatly benefit from efficient tokenization. However, they still mostly utilize the classical Byte-Pair Encoding (BPE) algorithm, a simple and reliable method. BPE has been shown to cause such issues as undertrained tokens and sub-optimal compression that may affect the downstream performance. We introduce PickyBPE, a modified BPE algorithm that carries out vocabulary refinement during tokenizer training by removing merges that leave intermediate "junk" tokens. Our method improves vocabulary efficiency, eliminates under-trained tokens, and does not compromise text compression. Our experiments show that this method either improves downstream performance or does not harm it.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f41391f7-6f0d-4c05-ab65-75d6aee611d6Cited by top-tier papers9
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine et al.EMNLP 2024 · 16 citations
- Sampling from Your Language Model One Byte at a TimeJonathan Hayase, Alisa Liu, Noah Smith, Sewoong OhICML 2026 · 9 citations
- Exploiting Vocabulary Frequency Imbalance in Language Model Pre-trainingWoojin Chung, Jeonghoon KimNeurIPS 2025 · 6 citations
- A Partition Cover Approach to TokenizationJia Peng Lim, Shawn Tan, Davin Choo, Hady W. LauwNeurIPS 2025 · 6 citations
- Lossless Vocabulary Reduction for Auto-Regressive Language ModelsDaiki Chijiwa, Taku Hasegawa, Kyosuke Nishida, Shin'ya Yamaguchi et al.ICLR 2026 · 3 citations
Builds on11
- Language Model Tokenizers Introduce Unfairness Between LanguagesAleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, Adel BibiNeurIPS 2023 · 301 citations
- OLMo: Accelerating the Science of Language ModelsDirk Groeneveld, Iz Beltagy, Evan Pete Walsh, Akshita Bhagia et al.ACL 2024 · 52 citations
- XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language ModelsDavis Liang, Hila Gonen, Yuning Mao, Rui Hou et al.EMNLP 2023 · 29 citations
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 17 citations
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine et al.EMNLP 2024 · 16 citations
Related papers
- Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language ModelsTomohiro Sawada, Kartik GoyalEMNLP 2025
- Scaffold-BPE: Enhancing Byte Pair Encoding for Large Language Models with Simple and Effective Scaffold Token RemovalHaoran Lian, Yizhe Xiong, Jianwei Niu, Shasha Mo et al.AAAI 2025
- Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in TokenizationNegar Foroutan, Clara Meister, Debjit Paul, Joel Niklaus et al.ACL 2026 · 14 citations
- Incremental BPE TokenizationShenghu Jiang, Ruihao GongICML 2026 · 12 citations
- From Where Words Come: Efficient Regularization of Code Tokenizers Through Source AttributionPavel Chizhov, Egor Bogomolov, Ivan P. YamshchikovACL 2026
