BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
Pavel Chizhov, Catherine Arnett, Elizaveta Korotkova, Ivan P. Yamshchikov
摘要
Language models can greatly benefit from efficient tokenization. However, they still mostly utilize the classical Byte-Pair Encoding (BPE) algorithm, a simple and reliable method. BPE has been shown to cause such issues as undertrained tokens and sub-optimal compression that may affect the downstream performance. We introduce PickyBPE, a modified BPE algorithm that carries out vocabulary refinement during tokenizer training by removing merges that leave intermediate "junk" tokens. Our method improves vocabulary efficiency, eliminates under-trained tokens, and does not compromise text compression. Our experiments show that this method either improves downstream performance or does not harm it.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine 等EMNLP 2024 · 被引用 16 次
- Sampling from Your Language Model One Byte at a TimeJonathan Hayase, Alisa Liu, Noah Smith, Sewoong OhICML 2026 · 被引用 9 次
- Exploiting Vocabulary Frequency Imbalance in Language Model Pre-trainingWoojin Chung, Jeonghoon KimNeurIPS 2025 · 被引用 6 次
- A Partition Cover Approach to TokenizationJia Peng Lim, Shawn Tan, Davin Choo, Hady W. LauwNeurIPS 2025 · 被引用 6 次
- Lossless Vocabulary Reduction for Auto-Regressive Language ModelsDaiki Chijiwa, Taku Hasegawa, Kyosuke Nishida, Shin'ya Yamaguchi 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper11
- Language Model Tokenizers Introduce Unfairness Between LanguagesAleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, Adel BibiNeurIPS 2023 · 被引用 301 次
- OLMo: Accelerating the Science of Language ModelsDirk Groeneveld, Iz Beltagy, Evan Pete Walsh, Akshita Bhagia 等ACL 2024 · 被引用 52 次
- XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language ModelsDavis Liang, Hila Gonen, Yuning Mao, Rui Hou 等EMNLP 2023 · 被引用 29 次
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 被引用 17 次
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine 等EMNLP 2024 · 被引用 16 次
相关 Paper
- Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language ModelsTomohiro Sawada, Kartik GoyalEMNLP 2025
- Scaffold-BPE: Enhancing Byte Pair Encoding for Large Language Models with Simple and Effective Scaffold Token RemovalHaoran Lian, Yizhe Xiong, Jianwei Niu, Shasha Mo 等AAAI 2025
- Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in TokenizationNegar Foroutan, Clara Meister, Debjit Paul, Joel Niklaus 等ACL 2026 · 被引用 14 次
- Incremental BPE TokenizationShenghu Jiang, Ruihao GongICML 2026 · 被引用 12 次
- From Where Words Come: Efficient Regularization of Code Tokenizers Through Source AttributionPavel Chizhov, Egor Bogomolov, Ivan P. YamshchikovACL 2026
