Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective Detection
Yuxi Li, Yi Liu, Gelei Deng, Ying Zhang, Wenjia Song, Ling Shi, Kailong Wang, Yuekang Li, Yang Liu, Haoyu Wang
Abstract
With the expanding application of Large Language Models (LLMs) in various domains, it becomes imperative to comprehensively investigate their unforeseen behaviors and consequent outcomes. In this study, we introduce and systematically explore the phenomenon of "glitch tokens", which are anomalous tokens produced by established tokenizers and could potentially compromise the models' quality of response. Specifically, we experiment on seven top popular LLMs utilizing three distinct tokenizers and involving a totally of 182,517 tokens. We present categorizations of the identified glitch tokens and symptoms exhibited by LLMs when interacting with glitch tokens. Based on our observation that glitch tokens tend to cluster in the embedding space, we propose GlitchHunter, a novel iterative clustering-based technique, for efficient glitch token detection. The evaluation shows that our approach notably outperforms three baseline methods on eight open-source LLMs. To the best of our knowledge, we present the first comprehensive study on glitch tokens. Our new detection further provides valuable insights into mitigating tokenization-related errors in LLMs. CCS Concepts: • Computing methodologies → Knowledge representation and reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 000b8187-ed58-48d9-ac4b-eb995c1dd1baCited by top-tier papers16
- PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration TestingGelei Deng, Yi Liu, Víctor Mayoral Vilches, Peng Liu et al.USENIX Security 2024 · 186 citations
- Token Embeddings Violate the Manifold HypothesisMichael Robinson, Sourya Dey, Tony ChiangNeurIPS 2025 · 19 citations
- Semantic-Enhanced Indirect Call Analysis with Large Language ModelsBaijun Cheng, Cen Zhang, Kailong Wang, Ling Shi et al.ASE 2024 · 4 citations
- Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak AttacksShide Zhou, Tianlin Li, Kailong Wang, Yihao Huang et al.ICSE 2025 · 3 citations
- GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language ModelsZhibo Zhang, Wuxia Bai, Yuxi Li, Mark Huasong Meng et al.ASE 2024 · 2 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang et al.ICLR 2023 · 295 citations
- An Empirical Study on Fine-Tuning Large Language Models of Code for Automated Program RepairKai Huang, Xiangxin Meng, Jian Zhang, Yang Liu et al.ASE 2023 · 91 citations
- NNSmith: Generating Diverse and Valid Test Cases for Deep Learning CompilersJiawei Liu, Jinkun Lin, Fabian Ruffy, Cheng Tan et al.ASPLOS 2023 · 90 citations
- BiasAsker: Measuring the Bias in Conversational AI SystemYuxuan Wan, Wenxuan Wang, Pinjia He, Jiazhen Gu et al.FSE 2023 · 50 citations
Related papers
- GlitchMiner: Mining Glitch Tokens in Large Language Models via Gradient-based Discrete OptimizationZihui Wu, Haichang Gao, Ping Wang, Shudong Zhang et al.AAAI 2026 · 1 citation
- One Bad Token Spoils the Barrel: Assessment, Detection, and Remediation of Glitch Tokens in Large Language ModelsKunsheng Tang, Peigui Qi, Yide Song, Wenbo Zhou et al.USENIX Security 2026
- Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsSander Land, Max BartoloEMNLP 2024 · 4 citations
- GlitchCleaner: Lightweight Glitch Tokens Repairing by Lossless Gated LoRA in Large Language ModelsYibo Fan, Jingru Li, Huan LiAAAI 2026
- Sticking to the Mean: Detecting Sticky Tokens in Text Embedding ModelsKexin Chen, Dongxia Wang, Yi Liu, Haonan Zhang et al.ACL 2025
