Tokenisation is NP-Complete
Philip Whittington, Gregor Bachmann, Tiago Pimentel
2025年份
6被引次数
3顶会引用
摘要
In this work, we prove the NP-completeness of two variants of tokenisation, defined as the problem of compressing a dataset to at most symbols by either finding a vocabulary directly (direct tokenisation), or selecting a sequence of merge operations (bottom-up tokenisation).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A Partition Cover Approach to TokenizationJia Peng Lim, Shawn Tan, Davin Choo, Hady W. LauwNeurIPS 2025 · 被引用 6 次
- so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMsSriharsh Bhyravajjula, Melanie Walsh, Anna Preus, Maria AntoniakEMNLP 2025 · 被引用 1 次
- Causal Estimation of Tokenisation BiasPietro Lesci, Clara Meister, Thomas Hofmann, Andreas Vlachos 等ACL 2025
它引用的顶会 Paper9
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- MEGABYTE: Predicting Million-byte Sequences with Multiscale TransformersLili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan 等NeurIPS 2023 · 被引用 197 次
- Do All Languages Cost the Same? Tokenization in the Era of Commercial Language ModelsOrevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai 等EMNLP 2023 · 被引用 24 次
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine 等EMNLP 2024 · 被引用 16 次
- Tokenization and the Noiseless ChannelVilém Zouhar, Clara Meister, Juan Luis Gastaldi, Li Du 等ACL 2023 · 被引用 10 次
相关 Paper
- Tokenisation over Bounded Alphabets is HardVioleta Kastreva, Philip Whittington, Dennis Komm, Tiago PimentelICLR 2026 · 被引用 6 次
- Efficient Algorithms for the Uniform Tokenization ProblemAngela W. Li, Konstantinos MamourasOOPSLA 2025 · 被引用 3 次
- Static Analysis for Efficient Streaming TokenizationAngela W. Li, Yudi Yang, Konstantinos MamourasASPLOS 2026 · 被引用 1 次
- Explaining and Mitigating Crosslingual Tokenizer InequitiesCatherine Arnett, Tyler A. Chang, Stella Biderman, Benjamin BergenNeurIPS 2025 · 被引用 9 次
- BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer TrainingPavel Chizhov, Catherine Arnett, Elizaveta Korotkova, Ivan P. YamshchikovEMNLP 2024 · 被引用 1 次
