USENIX Security2026Top-tier venue
One Bad Token Spoils the Barrel: Assessment, Detection, and Remediation of Glitch Tokens in Large Language Models
Kunsheng Tang, Peigui Qi, Yide Song, Wenbo Zhou, Zhicong Huang, Qing Guo, Tianwei Zhang, Weiming Zhang, Nenghai Yu, Jie Zhang
Abstract
Large Language Models (LLMs) have shown remarkable capabilities across numerous applications. However, the recent emergence of "glitch tokens," referring to tokens that cause unpredictable and erroneous model behaviors, poses significant reliability and security challenges. Despite prior investigations of glitch tokens, critical gaps remain, including insufficient safety assessments, limited detection coverage, and the absence of effective remediation strategies. To address these challenges, we first demonstrate that glitch tokens can readily bypass safety mechanisms and elicit unsafe outputs across LLMs through straightforward exploitation methods. More critically, we reveal that these tokens exhibit cross-model transferability, inducing safety risks across model families, tokenizers, commercial and moderation systems, underscoring their pervasive security implications. We then propose GlitchQuiz, a red-teaming framework inspired by human language acquisition, to comprehensively detect glitch tokens, surpassing existing detection methods. Finally, we develop GlitchEdit, a training-free embedding-layer editing approach that effectively remediates glitch tokens, reducing unsafe behaviors with average unsafe rates decreasing from 96.36% to 2.87% across evaluated LLMs and maintaining overall performance. Our findings have been responsibly disclosed to nine affected leading LLM providers, including OpenAI, Anthropic, and others, to help foster safer AI ecosystems.
Warning: This paper contains examples of unsafe content generated by large language models, including violence, discrimination, and other content that may be disturbing or offensive to some readers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6145a118-73ce-4553-9c33-0fbd18ee29f1Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li et al.ICLR 2024 · 419 citations
- Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and ReconstructionTong Liu, Yingjie Zhang, Zhe Zhao, Yinpeng Dong et al.USENIX Security 2024 · 121 citations
Related papers
- GlitchMiner: Mining Glitch Tokens in Large Language Models via Gradient-based Discrete OptimizationZihui Wu, Haichang Gao, Ping Wang, Shudong Zhang et al.AAAI 2026 · 1 citation
- GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language ModelsZhibo Zhang, Wuxia Bai, Yuxi Li, Mark Huasong Meng et al.ASE 2024 · 2 citations
- Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective DetectionYuxi Li, Yi Liu, Gelei Deng, Ying Zhang et al.FSE 2024 · 12 citations
- GlitchCleaner: Lightweight Glitch Tokens Repairing by Lossless Gated LoRA in Large Language ModelsYibo Fan, Jingru Li, Huan LiAAAI 2026
- Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsSander Land, Max BartoloEMNLP 2024 · 4 citations
