ENCHTABLE: Unified Safety Alignment Transfer in Fine-Tuned Large Language Models
Jialin Wu, Kecen Li, Zhicong Huang, Xinfeng Li, XiaoFeng Wang, Cheng Hong
摘要
Nowadays, many machine learning models are finetuned from large language models (LLMs) to achieve high performance in specialized domains such as code generation, biomedical analysis, and mathematical problem solving. However, researchers have shown that such fine-tuning process often introduces a critical vulnerability: the systematic degradation of safety alignment, which undermines ethical guidelines and increases the risk of harmful outputs. Addressing this challenge, we introduce ENCHTABLE, a novel and unified framework designed to transfer and maintain safety alignment in downstream LLMs without requiring extensive retraining. ENCHTABLE leverages a Neural Tangent Kernel (NTK)based safety vector distillation method to decouple safety constraints from task-specific reasoning, ensuring compatibility across diverse model architectures and sizes. Additionally, our interference-aware merging technique effectively balances the trade-offs between safety and utility, minimizing performance compromises across various task domains.
We have implemented a fully functional prototype of ENCHTABLE on three different task domains and three distinct LLM architectures, and evaluated its performance through extensive experiments on eleven diverse datasets, assessing both downstream utility and model safety. Our evaluations include assessments of LLMs from different vendors, demonstrating the generalization capability of ENCHTABLE. Furthermore, ENCHTABLE exhibits robust resistance to both static and dynamic jailbreaking attacks, outperforming vendor-released safety models in mitigating adversarially designed prompts. Comparative analyses with six parameter modification methods and two inference-time alignment baselines reveal that ENCHTABLE achieves significantly lower unsafe rate and higher utility score and universal applicability across different task domains. Additionally, we validate that ENCHTABLE can be seamlessly integrated into various deployment pipelines without significant overhead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- AgentAuditor: Human-level Safety and Security Evaluation for LLM AgentsHanjun Luo, Shenyu Dai, Chiming Ni, Xinfeng Li 等NeurIPS 2025 · 被引用 98 次
- Understanding and Preserving Safety in Fine-Tuned LLMsJiawen Zhang, Yangfan Hu, Kejia Chen, Lipeng He 等CCS 2026 · 被引用 7 次
- Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention SinkGuozhi Liu, Weiwei Lin, Tiansheng Huang, Ruichao Mo 等ICML 2026 · 被引用 5 次
它引用的顶会 Paper41
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
相关 Paper
- Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak AttacksYingjie Zhang, Tong Liu, Zhe Zhao, Guozhu Meng 等NDSS 2026 · 被引用 5 次
- SAFT: Safety-Preserving Adaptation via Fine-Tuning Transfer for Large Language ModelsZhiwen Ruan, Yan Yang, Zhuocheng Liang, Yun Chen 等KDD 2026
- Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning DatasetsLei Hsiung, Tianyu Pang, Yung-Chen Tang, Linyue Song 等ACL 2026 · 被引用 22 次
- AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety BasinShuo Yang, Qihui Zhang, Yuyang Liu, Yue Huang 等AAAI 2026 · 被引用 19 次
- Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template RegionChak Tou Leong, Qingyu Yin, Jian Wang, Wenjie LiACL 2025
