Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
Zhixin Xie, Xurui Song, Jun Luo
摘要
Despite substantial efforts in safety alignment, recent research indicates that Large Language Models (LLMs) remain highly susceptible to jailbreak attacks. Among these attacks, finetuning-based ones that compromise LLMs' safety alignment via fine-tuning stand out due to its stable jailbreak performance. In particular, a recent study indicates that fine-tuning with as few as 10 harmful question-answer (QA) pairs can lead to successful jailbreaking across various harmful questions. However, such malicious fine-tuning attacks are readily detectable and hence thwarted by moderation models. In this paper, we demonstrate that LLMs can be jailbroken by fine-tuning with only 10 benign QA pairs; our attack exploits the increased sensitivity of LLMs to fine-tuning data after being overfitted. Specifically, our fine-tuning process starts with overfitting an LLM via fine-tuning with benign QA pairs involving identical refusal answers. Further fine-tuning is then performed with standard benign answers, causing the overfitted LLM to forget the refusal attitude and thus provide compliant answers regardless of the harmfulness of a question. We implement our attack on the ten LLMs and compare it with five existing baselines. Experiments demonstrate that our method achieves significant advantages in both attack effectiveness and attack stealth. Our findings expose previously unreported security vulnerabilities in current LLMs and provide a new perspective on understanding how LLMs' security is compromised, even with benign fine-tuning. Our code is available at https://github.com/ZHIXINXIE/ ten_benign.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Self-Destructive Language ModelsYuhui Wang, Rongyi Zhu, Ting WangICLR 2026 · 被引用 14 次
- Where to Start Alignment? Diffusion Large Language Model May Demand a Distinct PositionZhixin Xie, Xurui Song, Jun LuoAAAI 2026 · 被引用 7 次
- Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update AlignmentYavuz Faruk Bakman, Duygu Nur Yaldiz, Eleni Triantafillou, Peter Kairouz 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
相关 Paper
- Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning DatasetsLei Hsiung, Tianyu Pang, Yung-Chen Tang, Linyue Song 等ACL 2026 · 被引用 22 次
- Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak AttacksYingjie Zhang, Tong Liu, Zhe Zhao, Guozhu Meng 等NDSS 2026 · 被引用 5 次
- Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?Sravanti Addepalli, Yerram Varun, Arun Suggala, Karthikeyan Shanmugam 等ICLR 2025
- Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language ModelsYongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang 等ICML 2024 · 被引用 140 次
- Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-CheckChentao Cao, Xiaojun Xu, Bo Han, Hang LiICLR 2026 · 被引用 5 次
