Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks
Yingjie Zhang, Tong Liu, Zhe Zhao, Guozhu Meng, Kai Chen
Abstract
Large Language Models (LLMs) remain vulnerable to jailbreak attacks that exploit adversarial prompts to circumvent safety measures. Current safety fine-tuning approaches face two critical limitations. First, they often fail to strike a balance between security and utility, where stronger safety measures tend to over-reject harmless user requests. Second, they frequently miss malicious intent concealed within seemingly benign tasks, leaving models exposed to exploitation. Our work identifies a fundamental cause of these issues: during response generation, an LLM's capacity to differentiate harmful from safe outputs deteriorates. Experimental evidence confirms this, revealing that the separability between hidden states for safe and harmful responses diminishes as generation progresses. This weakening discrimination forces models to make compliance judgments earlier in the generation process, restricting their ability to recognize developing harmful intent and contributing to both aforementioned failures. To mitigate this vulnerability, we introduce DEEPALIGN -an inherent defense framework that enhances the safety of LLMs. By applying contrastive hidden-state steering at the midpoint of response generation, DEEPALIGN amplifies the separation between harmful and benign hidden states, enabling continuous intrinsic toxicity detection and intervention throughout the generation process. Moreover, it facilitates contextually appropriate safe responses to harmful queries, thereby expanding the feasible space of safe responses. Evaluations demonstrate DEEPALIGN's efficacy. Across diverse LLMs spanning varying architectures and scales, it reduced attack success rates of nine distinct jailbreak attacks to near-zero or minimal. Crucially, it preserved model capability while reducing over-refusal. Models equipped with DEEPALIGN exhibited up to 3.5% lower error rates in rejecting challenging benign queries and maintained standard task performance with less than 1% decline. This marks a substantial advance in the safety-utility Pareto frontier. Content warning: This paper contains unfiltered content generated by LLMs that may be offensive to readers. * Corresponding authors and fuzz testing [36] , [56] . However, the increasing prevalence of LLMs has brought critical security concerns to the forefront. A notable issue is their susceptibility to jailbreak attacks [64], [43], [29], [48] , where crafted prompts bypass safeguards to generate harmful outputs. These pose significant risks, from societal harm to cybersecurity threats like remote code execution (RCE) [28] . Recent cases show real-world dangers, such as malware using public LLM APIs to dynamically create attack payloads that bypass detection [7] . Current defense primarily adopts two paradigms: external security guardrails and endogenous safeguards. While API providers offer guardrails, they face accuracy-recall tradeoffs. This challenge is particularly acute for smaller organizations and individual developers relying on platforms like Hugging Face, who lack access to enterprise-grade guardrails. This highlights the need for stronger endogenous safeguards. Endogenous safeguards are inherently limited because safety alignment captures human preferences-something pretraining alone cannot capture. Current approaches address this by fine-tuning models on malicious queries and refusal responses, teaching them to reject harmful inputs [49] . However, these safeguards remain operationally brittle. Adversaries exploit semantic ambiguities and generation dynamics to craft inputs that bypass alignment efforts, inducing models to propagate harmful content. This reveals two interconnected challenges: • Intent Disambiguation in Adversarial Contexts. Malicious intent is often artfully embedded within semantically complex prompts or benign tasks, confounding reliable identification of malicious intent. • The Security-Utility Pareto Trade-off. Security enhancements frequently increase refusal rates on benign prompts, degrading utility and user experience-a fundamental constraint on robust LLM deployment. Dynamic Degradation in Discriminative Capacity: A Core Vulnerability Underpinning Safety-Utility Trade-offs. Our research reveals a critical, previously overlooked vulnerability: during response generation, the model's inherent capability to distinguish between benign and harmful token sequences progressively degrades. This is reflected in a measurable phenomenon: as the model generates more harmful response
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e96ac976-20c1-4ff4-b112-46689c1abb7dCited by top-tier papers1
Ask how each one uses itBuilds on21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
Related papers
- Safety Alignment Should be Made More Than Just a Few Tokens DeepXiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma et al.ICLR 2025
- Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning DatasetsLei Hsiung, Tianyu Pang, Yung-Chen Tang, Linyue Song et al.ACL 2026 · 22 citations
- JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and ManipulationShenyi Zhang, Yuchen Zhai, Keyan Guo, Hongxin Hu et al.USENIX Security 2025
- Look Before You Leap: Enhance Attention and Vigilance Regarding Harmful Content with GuidelineLLMShaoqing Zhang, Zhuosheng Zhang, Kehai Chen, Rongxiang Weng et al.AAAI 2025 · 1 citation
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming VulnerabilityShojiro Yamabe, Jun SakumaICLR 2026 · 9 citations
