Submodular Optimization for Minimal Augmentation in Robust Language Model Alignment
CHING-CHIA KAO, Chia-Mu Yu, Chun-Shien Lu, Chu-song Chen
Abstract
Safety alignment of large language models is fragile: even small fine-tuning perturbations elastically revert behaviors toward those of the pre-training, with degradation inversely proportional to the size of the alignment set. We ask how to achieve safety alignment with minimal augmentation. To this end, we model augmentation as a set of group actions on sequences and formalize robustness gains as a normalized, monotone submodular function over transformations. We then leverage submodular optimization to select minimal augmentations that provably improve robustness. Experiments confirm that our approach efficiently restores safety alignment while minimizing the overhead of augmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
Related papers
- AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety BasinShuo Yang, Qihui Zhang, Yuyang Liu, Yue Huang et al.AAAI 2026 · 19 citations
- Safety at One Shot: Patching Fine-Tuned LLMs with A Single InstanceJiawen Zhang, Lipeng He, Kejia Chen, Jian Lou et al.ICLR 2026 · 10 citations
- Safety Depth in Large Language Models: A Markov Chain PerspectiveChing-Chia Kao, Chia-Mu Yu, Chun-Shien Lu, Chu-Song ChenNeurIPS 2025 · 2 citations
- Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case StudyKaustubh Ponkshe, Shaan Shah, Raghav Singhal, Praneeth VepakommaICLR 2026 · 9 citations
- Multilingual Safety Alignment Via Sparse Weight EditingJiaming Liang, Zhaoxin Wang, Handing WangICML 2026 · 3 citations
