Accelerating String-Heavy Queries with LLM Token Tables
Tobias Schmidt, Nicolas Schmitt, Thomas Neumann, Andreas Kipf
Abstract
Strings are the most common data type in modern database systems, yet they are often treated as an afterthought in high-performance data formats. While numerical data benefits from specialized, lightweight compression schemes, text is typically handled by general-purpose algorithms such as Zstd, LZ4, or Snappy, which require full-block decompression before processing. In this paper, we explore the potential of repurposing Large Language Model (LLM) tokenizers as a lightweight string compression scheme for databases, similar to FSST, but with a global token table shared across all tables and columns. Operators such as joins and aggregations can exploit this consistent encoding to defer decompression and process encoded values directly. We implement a global token table based on GPT-4's tokenizer in Umbra and demonstrate execution time improvements of up to 2× on string-heavy workloads, while reducing storage and memory consumption by up to 1.65×. Tokenizers integrate well with other compression algorithms, such as FSST, OnPair, or Zstd, while maintaining good compression ratios and high decompression throughput exceeding 6 GB/s on a single CPU core.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7785ad90-67e8-46d1-b5a7-a08ae5980766Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Neural Machine Translation with Byte-Level SubwordsChanghan Wang, Kyunghyun Cho, Jiatao GuAAAI 2020 · 213 citations
- Chimp: Efficient Lossless Floating Point Compression for Time Series DatabasesPanagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2022 · 76 citations
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo et al.VLDB 2024 · 59 citations
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 47 citations
Related papers
- FSST: Fast Random Access String CompressionPeter Boncz, Thomas Neumann, Viktor LeisVLDB 2020
- zip2zip: Inference-Time Adaptive Tokenization via Online CompressionSaibo Geng, Nathan Ranchin, Yunzhen Yao, Maxime Peyrard et al.NeurIPS 2025 · 5 citations
- UniGist: Towards General and Hardware-aligned Sequence-level Long Context CompressionChenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li et al.NeurIPS 2025 · 10 citations
- Encoding Spreadsheets for Large Language ModelsHaoyu Dong, Jianbo Zhao, Yuzhang Tian, Junyu Xiong et al.EMNLP 2024 · 4 citations
- Lexico: Extreme KV Cache Compression via Sparse Coding over Universal DictionariesJunhyuck Kim, Jongho Park, Jaewoong Cho, Dimitris PapailiopoulosICML 2025
