Inference-Time Personalized Safety Control via Paired Difference-in-Means Intervention
Tran Huynh, Ruoxi Jia
摘要
Safety preferences are inherently subjective, yet current LLM safety alignment methods often impose universal standards that fail to account for individual sensitivities. In this work, we propose an efficient, training-free method for personalized safety control via inference-time activation intervention. Our approach steers internal representations to suppress user-specific undesired content while preserving model utility. We systematically evaluate three strategies for estimating intervention directions: Instance-Level Contrast Shift (ILCS), Unpaired Mean Shift (UMS), and our primary method, Paired Contrast Mean Shift (PCMS). We provide theoretical insights into each approach and highlight the advantages of PCMS. Empirical results across diverse open-weight models demonstrate that our method effectively reduces undesired content in line with individual preferences, with minimal impact on helpfulness-enabling more adaptive and user-aligned LLM behavior. Physical Effects: * Pain or discomfort * Injuries (e.g., bruises, cuts) * Fatigue or exhaustion * Sleep disturbances * Changes in appetite or eating habits Mental and Emotional Effects: * Anxiety or fear * Depression or low self-esteem * Difficulty concentrating or making decisions * Feelings of shame, guilt, or embarrassment * Loss of trust in others or oneself * Difficulty forming relationships or maintaining friendships * Increased stress levels * Hypervigilance or exaggerated startle response
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 被引用 337 次
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya 等NeurIPS 2023 · 被引用 295 次
相关 Paper
- Who's asking? User personas and the mechanics of latent misalignmentAsma Ghandeharioun, Ann Yuan, Marius Guerard, Emily Reif 等NeurIPS 2024 · 被引用 44 次
- The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test AwarenessSahar Abdelnabi, Ahmed SalemNeurIPS 2025 · 被引用 28 次
- A Simple Yet Effective Method for Non-Refusing Context Relevant Fine-grained Safety Steering in LLMsShaona Ghosh, Amrita Bhattacharjee, Yftah Ziser, Christopher ParisienEMNLP 2025
- SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender SystemsHaochang Hao, Yifan Xu, Xinzhuo Li, Yingqiang Ge 等KDD 2026 · 被引用 1 次
- Steering Llama 2 via Contrastive Activation AdditionNina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong 等ACL 2024
