One Bad Token Spoils the Barrel: Assessment, Detection, and Remediation of Glitch Tokens in Large Language Models
Kunsheng Tang, Peigui Qi, Yide Song, Wenbo Zhou, Zhicong Huang, Qing Guo, Tianwei Zhang, Weiming Zhang, Nenghai Yu, Jie Zhang
摘要
Large Language Models (LLMs) have shown remarkable capabilities across numerous applications. However, the recent emergence of "glitch tokens," referring to tokens that cause unpredictable and erroneous model behaviors, poses significant reliability and security challenges. Despite prior investigations of glitch tokens, critical gaps remain, including insufficient safety assessments, limited detection coverage, and the absence of effective remediation strategies. To address these challenges, we first demonstrate that glitch tokens can readily bypass safety mechanisms and elicit unsafe outputs across LLMs through straightforward exploitation methods. More critically, we reveal that these tokens exhibit cross-model transferability, inducing safety risks across model families, tokenizers, commercial and moderation systems, underscoring their pervasive security implications. We then propose GlitchQuiz, a red-teaming framework inspired by human language acquisition, to comprehensively detect glitch tokens, surpassing existing detection methods. Finally, we develop GlitchEdit, a training-free embedding-layer editing approach that effectively remediates glitch tokens, reducing unsafe behaviors with average unsafe rates decreasing from 96.36% to 2.87% across evaluated LLMs and maintaining overall performance. Our findings have been responsibly disclosed to nine affected leading LLM providers, including OpenAI, Anthropic, and others, to help foster safer AI ecosystems.
Warning: This paper contains examples of unsafe content generated by large language models, including violence, discrimination, and other content that may be disturbing or offensive to some readers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li 等ICLR 2024 · 被引用 419 次
- Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and ReconstructionTong Liu, Yingjie Zhang, Zhe Zhao, Yinpeng Dong 等USENIX Security 2024 · 被引用 121 次
相关 Paper
- GlitchMiner: Mining Glitch Tokens in Large Language Models via Gradient-based Discrete OptimizationZihui Wu, Haichang Gao, Ping Wang, Shudong Zhang 等AAAI 2026 · 被引用 1 次
- GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language ModelsZhibo Zhang, Wuxia Bai, Yuxi Li, Mark Huasong Meng 等ASE 2024 · 被引用 2 次
- Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective DetectionYuxi Li, Yi Liu, Gelei Deng, Ying Zhang 等FSE 2024 · 被引用 12 次
- GlitchCleaner: Lightweight Glitch Tokens Repairing by Lossless Gated LoRA in Large Language ModelsYibo Fan, Jingru Li, Huan LiAAAI 2026
- Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsSander Land, Max BartoloEMNLP 2024 · 被引用 4 次
