Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
Jingyu Zhang, Ahmed Elgohary, Ahmed Magooda, Daniel Khashabi, Benjamin Van Durme
摘要
The current paradigm for safety alignment of large language models (LLMs) follows a one-size-fits-all approach: the model refuses to interact with any content deemed unsafe by the model provider. This approach lacks flexibility in the face of varying social norms across cultures and regions. In addition, users may have diverse safety needs, making a model with static safety standards too restrictive to be useful, as well as too costly to be re-aligned. We propose Controllable Safety Alignment (CoSA), a framework designed to adapt models to diverse safety requirements without re-training. Instead of aligning a fixed model, we align models to follow safety configs-free-form natural language descriptions of the desired safety behaviors-that are provided as part of the system prompt. To adjust model safety behavior, authorized users only need to modify such safety configs at inference time. To enable that, we propose CoSAlign, a data-centric method for aligning LLMs to easily adapt to diverse safety configs. Furthermore, we devise a novel controllability evaluation protocol that considers both helpfulness and configured safety, summarizing them into CoSA-Score, and construct CoSApien, a human-authored benchmark that consists of real-world LLM use cases with diverse safety requirements and corresponding evaluation prompts. We show that CoSAlign leads to substantial gains of controllability over strong baselines including in-context alignment. Our framework encourages better representation and adaptation to pluralistic human values in LLMs, and thereby increasing their practicality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- DynaGuard: A Dynamic Guardian Model With User-Defined PoliciesMonte Hoover, Vatsal Baherwani, Neel Jain, Khalid Saifullah 等ICLR 2026 · 被引用 18 次
- The Alignment Waltz: Jointly Training Agents to Collaborate for SafetyJingyu Zhang, Haozhu Wang, Eric Michael Smith, Sid Wang 等ICLR 2026 · 被引用 11 次
- PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI HarmJingjing Li, Joel Mire, Eve Fleisig, Valentina Pyatkin 等ICLR 2026 · 被引用 6 次
- COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMsDasol Choi, DongGeon Lee, Brigitta Jesica Kartono, Helena Berndt 等ACL 2026 · 被引用 3 次
- PICACO: Pluralistic In-Context Value Alignment via Total Correlation OptimizationHan Jiang, Dongyao Zhu, Xiaoyuan Yi, Ziang Xiao 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper27
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
相关 Paper
- Configurable Reward Model for Balanced Safety AlignmentZhengping Jiang, Mehran Khodabandeh, Akash Bharadwaj, Manik Bhandari 等ICML 2026
- AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement LearningYi Zhang, An Zhang, XiuYu Zhang, Leheng Sheng 等ICLR 2026 · 被引用 15 次
- Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuningGuoli Wang, Haonan Shi, Tu Ouyang, An WangKDD 2026 · 被引用 5 次
- Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case StudyKaustubh Ponkshe, Shaan Shah, Raghav Singhal, Praneeth VepakommaICLR 2026 · 被引用 9 次
- Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained OptimizationTuan Nguyen, Long Tran-ThanhICML 2026
