Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection
Zhipeng Wei, Yuqi Liu, N. Benjamin Erichson
Abstract
Jailbreaking techniques trick Large Language Models (LLMs) into producing restricted output, posing a potential threat. One line of defense is to use another LLM as a Judge to evaluate the harmfulness of generated text. However, we reveal that these Judge LLMs are vulnerable to token segmentation bias, an issue that arises when delimiters alter the tokenization process, splitting words into smaller sub-tokens. This alters the embeddings of the entire sequence, reducing detection accuracy and allowing harmful content to be misclassified as safe. In this paper, we introduce Emoji Attack, a novel strategy that amplifies existing jailbreak prompts by exploiting token segmentation bias. Our method leverages in-context learning to systematically insert emojis into text before it is evaluated by a Judge LLM, inducing embedding distortions that significantly lower the likelihood of detecting unsafe content. Unlike traditional delimiters, emojis also introduce semantic ambiguity, making them particularly effective in this attack. Through experiments on state-of-the-art Judge LLMs, we demonstrate that Emoji Attack substantially reduces the unsafe prediction rate, bypassing existing safeguards.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cd386dd-779b-4ea4-94dd-9af73293d66fCited by top-tier papers8
- Adaptive Attacks on Trusted Monitors Subvert AI Control ProtocolsMikhail Terekhov, Alexander Panfilov, Daniil Dzenhaliou, Caglar Gulcehre et al.ICLR 2026 · 26 citations
- Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen AttacksMahavir Dabas, Tran Huynh, Nikhil Reddy Billa, Jiachen T. Wang et al.ICLR 2026 · 6 citations
- AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning ModelsJiacheng Liang, Tanqiu Jiang, Yuhui Wang, Rongyi Zhu et al.ACL 2026 · 3 citations
- False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language ModelsWeipeng Jiang, Xiaoyu Zhang, Juan Zhai, Shiqing Ma et al.ACL 2026 · 1 citation
- Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM JudgesXIANGLIN YANG, Bryan Hooi, Gelei Deng, Tianwei Zhang et al.ICML 2026
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
Related papers
- Efficient LLM-Jailbreaking via Multimodal-LLM JailbreakHaoxuan Ji, Zheng Lin, Zhenxing Niu, Xinbo Gao et al.AAAI 2026 · 4 citations
- Towards Understanding Jailbreak Attacks in LLMs: A Representation Space AnalysisYuping Lin, Pengfei He, Han Xu, Yue Xing et al.EMNLP 2024 · 6 citations
- Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal BoundariesJiahao Yu, Haozheng Luo, Jerry Yao-Chieh Hu, Yan Chen et al.USENIX Security 2025
- Exploiting Task-Level Vulnerabilities: An Automatic Jailbreak Attack and Defense Benchmarking for LLMsLan Zhang, Xinben Gao, Liuyi Yao, Jinke Song et al.USENIX Security 2025
- When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' ToxicityShiyao Cui, Xijia Feng, Yingkang Wang, Junxiao Yang et al.AAAI 2026
