Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models
Tomohiro Sawada, Kartik Goyal
Abstract
Standard Byte-Pair Encoding (BPE) tokenization compresses text by pairing a learned token vocabulary with a detailed merge list. Recent work has shown that this merge list exposes a potential attack surface for extracting information about language model's training data. In this paper, we explore the downstream impact of BPE inference algorithms that do not rely on this merge list at all, and hence differ from the encoding process during BPE training. To address this question, we investigate two broad classes of BPE inference schemes that differ from BPE application during training: a) targeted deviation from merge-lists including random merge orders, and various corruptions of merge list involving deletion/truncation, and b) non-targeted BPE inference algorithms that do not depend on the merge list but focus on compressing the text either greedily or exactly. Extensive experiments across diverse language modeling tasks like accuracy-based QA benchmarks, machine translation, and open-ended generation reveal that while targeted deviation from the merge lists exhibits significant degradation in language model performance, the non-targeted merge-list-free inference algorithms result in minimal impact on downstream performance that is often much smaller than expected. These findings pave way for simpler and potentially more privacy-preserving tokenization schemes that do not catastrophically compromise model performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 56a40d97-43c8-4e4c-b915-742ff5fb5079Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- S2ORC: The Semantic Scholar Open Research CorpusKyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney et al.ACL 2020 · 424 citations
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 17 citations
- You should evaluate your language model on marginal likelihood over tokenisationsKris Cao, Laura RimellEMNLP 2021 · 6 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
Related papers
- BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer TrainingPavel Chizhov, Catherine Arnett, Elizaveta Korotkova, Ivan P. YamshchikovEMNLP 2024 · 1 citation
- A Partition Cover Approach to TokenizationJia Peng Lim, Shawn Tan, Davin Choo, Hady W. LauwNeurIPS 2025 · 6 citations
- Data Mixture Inference Attack: BPE Tokenizers Reveal Training Data CompositionsJonathan Hayase, Alisa Liu, Yejin Choi, Sewoong Oh et al.NeurIPS 2024 · 16 citations
- Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in TokenizationNegar Foroutan, Clara Meister, Debjit Paul, Joel Niklaus et al.ACL 2026 · 14 citations
- An Efficient Algorithm for Streaming BPE TokenizationKonstantinos Mamouras, Angela W. Li, Yudi YangPLDI 2026
