ACL2025

Tokenisation is NP-Complete

Philip Whittington, Gregor Bachmann, Tiago Pimentel

被引用 6 次

摘要

In this work, we prove the NP-completeness of two variants of tokenisation, defined as the problem of compressing a dataset to at most δ\delta symbols by either finding a vocabulary directly (direct tokenisation), or selecting a sequence of merge operations (bottom-up tokenisation).