ACL2026
SGT: Securing Open-Source LLMs Against Malicious Fine-tuning via Safety Guidance Trigger
Sunguk Shin, Fangzhao Wu, Byung-Jun Lee, Meeyoung Cha, Sungwon Park
Abstract
Open-weight large language models (LLMs) enable extensive customization but remain susceptible to post-release misuse via malicious fine-tuning. While existing defenses attempt to constrain parameter-space dynamics or mitigate harmful internal representations, malicious fine-tuning continues to erode these safeguards leaving the development of fundamental, persistent defenses for open-weight models an unresolved challenge. In this paper, we characterize a safety region for open-weight LLMs and propose Safety Guidance Trigger (SGT), a framework that preserves alignment by guiding finetuning toward the safety manifold. It has two stages: (1) optimizing a safety trigger to steer the base model outputs toward safe responses and ( 2 ) training the open-weight model to align its internal features with trigger-induced safety representations. We demonstrate that SGT substantially improves robustness against malicious fine-tuning, forcing adversaries to significantly increase data budgets to bypass safeguards. Our analysis further confirms that this approach anchors model representations within a safety region that remains resilient under adversarial attacks.