Lune

ACL2026顶会

SGT: Securing Open-Source LLMs Against Malicious Fine-tuning via Safety Guidance Trigger

Sunguk Shin, Fangzhao Wu, Byung-Jun Lee, Meeyoung Cha, Sungwon Park

2026年份

摘要

Open-weight large language models (LLMs) enable extensive customization but remain susceptible to post-release misuse via malicious fine-tuning. While existing defenses attempt to constrain parameter-space dynamics or mitigate harmful internal representations, malicious fine-tuning continues to erode these safeguards leaving the development of fundamental, persistent defenses for open-weight models an unresolved challenge. In this paper, we characterize a safety region for open-weight LLMs and propose Safety Guidance Trigger (SGT), a framework that preserves alignment by guiding finetuning toward the safety manifold. It has two stages: (1) optimizing a safety trigger to steer the base model outputs toward safe responses and ( 2 ) training the open-weight model to align its internal features with trigger-induced safety representations. We demonstrate that SGT substantially improves robustness against malicious fine-tuning, forcing adversaries to significantly increase data budgets to bypass safeguards. Our analysis further confirms that this approach anchors model representations within a safety region that remains resilient under adversarial attacks.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper19

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖