GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models
Zhibo Zhang, Wuxia Bai, Yuxi Li, Mark Huasong Meng, Kailong Wang, Ling Shi, Li Li, Jun Wang, Haoyu Wang
Abstract
Large language models (LLMs) have achieved unprecedented success in the field of natural language processing. However, the blackbox nature of their internal mechanisms has brought many concerns about their trustworthiness and interpretability. Recent research has discovered a class of abnormal tokens in the model's vocabulary space and named them "glitch tokens". Those tokens, once included in the input, may induce the model to produce incorrect, irrelevant, or even harmful results, drastically undermining the reliability and practicality of LLMs. In this work, we aim to enhance the understanding of glitch tokens and propose techniques for their detection and mitigation. We first reveal the characteristic features induced by glitch tokens on LLMs, which are evidenced by significant deviations in the distributions of attention patterns and dynamic information from intermediate model layers. Based on the insights, we develop GlitchProber, a tool for efficient glitch token detection and mitigation. GlitchProber utilizes small-scale sampling, principal component analysis for accelerated feature extraction, and a simple classifier for efficient vocabulary screening. Taking one step further, GlitchProber rectifies abnormal model intermediate layer values to mitigate the destructive effects of glitch tokens. Evaluated on * Co-first author with equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 547b8d4f-db9b-4ff3-96f7-2a762abf4e12Cited by top-tier papers5
- Drowzee: Metamorphic Testing for Fact-Conflicting Hallucination Detection in Large Language ModelsNingke Li, Yuekang Li, Yi Liu, Ling Shi et al.OOPSLA 2024 · 26 citations
- Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak AttacksShide Zhou, Tianlin Li, Kailong Wang, Yihao Huang et al.ICSE 2025 · 3 citations
- Sticking to the Mean: Detecting Sticky Tokens in Text Embedding ModelsKexin Chen, Dongxia Wang, Yi Liu, Haonan Zhang et al.ACL 2025
- One Bad Token Spoils the Barrel: Assessment, Detection, and Remediation of Glitch Tokens in Large Language ModelsKunsheng Tang, Peigui Qi, Yide Song, Wenbo Zhou et al.USENIX Security 2026
- GlitchCleaner: Lightweight Glitch Tokens Repairing by Lossless Gated LoRA in Large Language ModelsYibo Fan, Jingru Li, Huan LiAAAI 2026
Builds on6
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Pay Attention to MLPsHanxiao Liu, Zihang Dai, David R. So, Quoc V. LeNeurIPS 2021 · 912 citations
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song et al.ICLR 2021 · 122 citations
- Drowzee: Metamorphic Testing for Fact-Conflicting Hallucination Detection in Large Language ModelsNingke Li, Yuekang Li, Yi Liu, Ling Shi et al.OOPSLA 2024 · 26 citations
- Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective DetectionYuxi Li, Yi Liu, Gelei Deng, Ying Zhang et al.FSE 2024 · 12 citations
Related papers
- GlitchMiner: Mining Glitch Tokens in Large Language Models via Gradient-based Discrete OptimizationZihui Wu, Haichang Gao, Ping Wang, Shudong Zhang et al.AAAI 2026 · 1 citation
- Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsSander Land, Max BartoloEMNLP 2024 · 4 citations
- Interpreting the Repeated Token Phenomenon in Large Language ModelsItay Yona, Ilia Shumailov, Jamie Hayes, Yossi GandelsmanICML 2025
- Attention Sinks as Internal Signals for Hallucination Detection in Large Language ModelsJakub Binkowski, Kamil Adamczewski, Tomasz KajdanowiczICML 2026
- On Early Detection of Hallucinations in Factual Question AnsweringBen Snyder, Marius Moisescu, Muhammad Bilal ZafarKDD 2024 · 10 citations
