ENCHTABLE: Unified Safety Alignment Transfer in Fine-Tuned Large Language Models
Jialin Wu, Kecen Li, Zhicong Huang, Xinfeng Li, XiaoFeng Wang, Cheng Hong
Abstract
Nowadays, many machine learning models are finetuned from large language models (LLMs) to achieve high performance in specialized domains such as code generation, biomedical analysis, and mathematical problem solving. However, researchers have shown that such fine-tuning process often introduces a critical vulnerability: the systematic degradation of safety alignment, which undermines ethical guidelines and increases the risk of harmful outputs. Addressing this challenge, we introduce ENCHTABLE, a novel and unified framework designed to transfer and maintain safety alignment in downstream LLMs without requiring extensive retraining. ENCHTABLE leverages a Neural Tangent Kernel (NTK)based safety vector distillation method to decouple safety constraints from task-specific reasoning, ensuring compatibility across diverse model architectures and sizes. Additionally, our interference-aware merging technique effectively balances the trade-offs between safety and utility, minimizing performance compromises across various task domains.
We have implemented a fully functional prototype of ENCHTABLE on three different task domains and three distinct LLM architectures, and evaluated its performance through extensive experiments on eleven diverse datasets, assessing both downstream utility and model safety. Our evaluations include assessments of LLMs from different vendors, demonstrating the generalization capability of ENCHTABLE. Furthermore, ENCHTABLE exhibits robust resistance to both static and dynamic jailbreaking attacks, outperforming vendor-released safety models in mitigating adversarially designed prompts. Comparative analyses with six parameter modification methods and two inference-time alignment baselines reveal that ENCHTABLE achieves significantly lower unsafe rate and higher utility score and universal applicability across different task domains. Additionally, we validate that ENCHTABLE can be seamlessly integrated into various deployment pipelines without significant overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8105b83f-9696-403a-af1d-013edfab987dCited by top-tier papers3
- AgentAuditor: Human-level Safety and Security Evaluation for LLM AgentsHanjun Luo, Shenyu Dai, Chiming Ni, Xinfeng Li et al.NeurIPS 2025 · 98 citations
- Understanding and Preserving Safety in Fine-Tuned LLMsJiawen Zhang, Yangfan Hu, Kejia Chen, Lipeng He et al.CCS 2026 · 7 citations
- Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention SinkGuozhi Liu, Weiwei Lin, Tiansheng Huang, Ruichao Mo et al.ICML 2026 · 5 citations
Builds on41
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
Related papers
- Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak AttacksYingjie Zhang, Tong Liu, Zhe Zhao, Guozhu Meng et al.NDSS 2026 · 5 citations
- SAFT: Safety-Preserving Adaptation via Fine-Tuning Transfer for Large Language ModelsZhiwen Ruan, Yan Yang, Zhuocheng Liang, Yun Chen et al.KDD 2026
- Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning DatasetsLei Hsiung, Tianyu Pang, Yung-Chen Tang, Linyue Song et al.ACL 2026 · 22 citations
- AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety BasinShuo Yang, Qihui Zhang, Yuyang Liu, Yue Huang et al.AAAI 2026 · 19 citations
- Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template RegionChak Tou Leong, Qingyu Yin, Jian Wang, Wenjie LiACL 2025
