SGT: Securing Open-Source LLMs Against Malicious Fine-tuning via Safety Guidance Trigger
Sunguk Shin, Fangzhao Wu, Byung-Jun Lee, Meeyoung Cha, Sungwon Park
Abstract
Open-weight large language models (LLMs) enable extensive customization but remain susceptible to post-release misuse via malicious fine-tuning. While existing defenses attempt to constrain parameter-space dynamics or mitigate harmful internal representations, malicious fine-tuning continues to erode these safeguards leaving the development of fundamental, persistent defenses for open-weight models an unresolved challenge. In this paper, we characterize a safety region for open-weight LLMs and propose Safety Guidance Trigger (SGT), a framework that preserves alignment by guiding finetuning toward the safety manifold. It has two stages: (1) optimizing a safety trigger to steer the base model outputs toward safe responses and ( 2 ) training the open-weight model to align its internal features with trigger-induced safety representations. We demonstrate that SGT substantially improves robustness against malicious fine-tuning, forcing adversaries to significantly increase data budgets to bypass safeguards. Our analysis further confirms that this approach anchors model representations within a safety region that remains resilient under adversarial attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2311d32a-b3f8-4186-b213-53d535a2d0c8Builds on19
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- Safe LoRA: The Silver Lining of Reducing Safety Risks when Finetuning Large Language ModelsChia-Yi Hsu, Yu-Lin Tsai, Chih-Hsun Lin, Pin-Yu Chen et al.NeurIPS 2024 · 165 citations
- Representation Noising: A Defence Mechanism Against Harmful FinetuningDomenic Rosati, Jan Wehner, Kai Williams, Lukasz Bartoszcze et al.NeurIPS 2024 · 107 citations
- Consistency Regularization for Certified Robustness of Smoothed ClassifiersJongheon Jeong, Jinwoo ShinNeurIPS 2020 · 103 citations
Related papers
- SDD: Self-Degraded Defense against Malicious Fine-tuningZixuan Chen, Weikai Lu, Xin Lin, Ziqian ZengACL 2025
- AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMsDebdeep Sanyal, Manodeep Ray, Murari MandalAAAI 2026 · 2 citations
- Tamper-Resistant Safeguards for Open-Weight LLMsRishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou et al.ICLR 2025
- AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety BasinShuo Yang, Qihui Zhang, Yuyang Liu, Yue Huang et al.AAAI 2026 · 19 citations
- Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language ModelsShengyun Peng, Pin-Yu Chen, Matthew Hull, Duen Horng ChauNeurIPS 2024 · 68 citations
