GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models
Zhibo Zhang, Wuxia Bai, Yuxi Li, Mark Huasong Meng, Kailong Wang, Ling Shi, Li Li, Jun Wang, Haoyu Wang
摘要
Large language models (LLMs) have achieved unprecedented success in the field of natural language processing. However, the blackbox nature of their internal mechanisms has brought many concerns about their trustworthiness and interpretability. Recent research has discovered a class of abnormal tokens in the model's vocabulary space and named them "glitch tokens". Those tokens, once included in the input, may induce the model to produce incorrect, irrelevant, or even harmful results, drastically undermining the reliability and practicality of LLMs. In this work, we aim to enhance the understanding of glitch tokens and propose techniques for their detection and mitigation. We first reveal the characteristic features induced by glitch tokens on LLMs, which are evidenced by significant deviations in the distributions of attention patterns and dynamic information from intermediate model layers. Based on the insights, we develop GlitchProber, a tool for efficient glitch token detection and mitigation. GlitchProber utilizes small-scale sampling, principal component analysis for accelerated feature extraction, and a simple classifier for efficient vocabulary screening. Taking one step further, GlitchProber rectifies abnormal model intermediate layer values to mitigate the destructive effects of glitch tokens. Evaluated on * Co-first author with equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Drowzee: Metamorphic Testing for Fact-Conflicting Hallucination Detection in Large Language ModelsNingke Li, Yuekang Li, Yi Liu, Ling Shi 等OOPSLA 2024 · 被引用 26 次
- Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak AttacksShide Zhou, Tianlin Li, Kailong Wang, Yihao Huang 等ICSE 2025 · 被引用 3 次
- Sticking to the Mean: Detecting Sticky Tokens in Text Embedding ModelsKexin Chen, Dongxia Wang, Yi Liu, Haonan Zhang 等ACL 2025
- One Bad Token Spoils the Barrel: Assessment, Detection, and Remediation of Glitch Tokens in Large Language ModelsKunsheng Tang, Peigui Qi, Yide Song, Wenbo Zhou 等USENIX Security 2026
- GlitchCleaner: Lightweight Glitch Tokens Repairing by Lossless Gated LoRA in Large Language ModelsYibo Fan, Jingru Li, Huan LiAAAI 2026
它引用的顶会 Paper6
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Pay Attention to MLPsHanxiao Liu, Zihang Dai, David R. So, Quoc V. LeNeurIPS 2021 · 被引用 912 次
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 等ICLR 2021 · 被引用 122 次
- Drowzee: Metamorphic Testing for Fact-Conflicting Hallucination Detection in Large Language ModelsNingke Li, Yuekang Li, Yi Liu, Ling Shi 等OOPSLA 2024 · 被引用 26 次
- Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective DetectionYuxi Li, Yi Liu, Gelei Deng, Ying Zhang 等FSE 2024 · 被引用 12 次
相关 Paper
- GlitchMiner: Mining Glitch Tokens in Large Language Models via Gradient-based Discrete OptimizationZihui Wu, Haichang Gao, Ping Wang, Shudong Zhang 等AAAI 2026 · 被引用 1 次
- Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsSander Land, Max BartoloEMNLP 2024 · 被引用 4 次
- Interpreting the Repeated Token Phenomenon in Large Language ModelsItay Yona, Ilia Shumailov, Jamie Hayes, Yossi GandelsmanICML 2025
- Attention Sinks as Internal Signals for Hallucination Detection in Large Language ModelsJakub Binkowski, Kamil Adamczewski, Tomasz KajdanowiczICML 2026
- On Early Detection of Hallucinations in Factual Question AnsweringBen Snyder, Marius Moisescu, Muhammad Bilal ZafarKDD 2024 · 被引用 10 次
