T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
Björn Deiseroth, Manuel Brack, Patrick Schramowski, Kristian Kersting, Samuel Weinbach
Abstract
Tokenizers are crucial for encoding information in Large Language Models, but their development has recently stagnated, and they contain inherent weaknesses. Major limitations include computational overhead, ineffective vocabulary use, and unnecessarily large embedding and head layers. Additionally, their performance is biased towards a reference corpus, leading to reduced effectiveness for underrepresented languages. To remedy these issues, we propose T-FREE which directly embeds words through sparse activation patterns over character triplets, and does not require a reference corpus. T-FREE inherently exploits morphological similarities and allows for strong compression of embedding layers. In our exhaustive experimental evaluation, we achieve competitive downstream performance with a parameter reduction of more than 85% on these layers. Further, T-FREE shows significant improvements in cross-lingual transfer learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64436da1-0b85-4417-abda-6de2aa0c8412Cited by top-tier papers3
- Scaling Embedding Layers in Language ModelsDa Yu, Edith Cohen, Badih Ghazi, Yangsibo Huang et al.NeurIPS 2025 · 21 citations
- From Bytes to Ideas: Language Modeling with Autoregressive U-NetsMathurin Videau, Badr Youbi Idrissi, Alessandro Ferreira Leite, Marc Schoenauer et al.NeurIPS 2025 · 14 citations
- Multidimensional Byte Pair Encoding: Shortened Sequences for Improved Visual Data GenerationTim Elsner, Paula Usinger, Julius Nehring-Wirxel, Gregor Kobsik et al.ICCV 2025
Builds on3
- Charformer: Fast Character Transformers via Gradient-based Subword TokenizationYi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta et al.ICLR 2022 · 198 citations
- MEGABYTE: Predicting Million-byte Sequences with Multiscale TransformersLili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan et al.NeurIPS 2023 · 197 citations
- HashFormers: Towards Vocabulary-independent Pre-trained TransformersHuiyin Xue, Nikolaos AletrasEMNLP 2022 · 2 citations
Related papers
- TokAlign: Efficient Vocabulary Adaptation via Token AlignmentChong Li, Jiajun Zhang, Chengqing ZongACL 2025 · 7 citations
- Unsupervised Tokenization LearningAnton Kolonin, Vignav RameshEMNLP 2022 · 3 citations
- Explaining and Mitigating Crosslingual Tokenizer InequitiesCatherine Arnett, Tyler A. Chang, Stella Biderman, Benjamin BergenNeurIPS 2025 · 9 citations
- Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language ModelsPit Neitemeier, Björn Deiseroth, Constantin Eichenberg, Lukas BallesICLR 2025
- PinTok: Tokenizers Deserve Dedicated Pinned CPU-Compute and MemorySean Choi, Myungheon Chin, Ernest RyuICML 2026
