Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach
Yuchen Wu, Edward Sun, Kaijie Zhu, Jianxun Lian, José Hernández-Orallo, Aylin Caliskan, Jindong Wang
Abstract
Large language models (LLMs) typically generate identical or similar responses for all users given the same prompt, posing serious safety risks in high-stakes applications where user vulnerabilities differ widely. Existing safety evaluations primarily rely on context-independent metrics - such as factuality, bias, or toxicity - overlooking the fact that the same response may carry divergent risks depending on the user's background or condition. We introduce personalized safety to fill this gap and present PENGUIN - a benchmark comprising 14,000 scenarios across seven sensitive domains with both context-rich and context-free variants. Evaluating six leading LLMs, we demonstrate that personalized user information significantly improves safety scores by 43.2%, confirming the effectiveness of personalization in safety alignment. However, not all context attributes contribute equally to safety enhancement. To address this, we develop RAISE - a training-free, two-stage agent framework that strategically acquires user-specific background. RAISE improves safety scores by up to 31.6% over six vanilla LLMs, while maintaining a low interaction cost of just 2.7 user queries on average. Our findings highlight the importance of selective information gathering in safety-critical domains and offer a practical solution for personalizing LLM responses without model retraining. This work establishes a foundation for safety research that adapts to individual user contexts rather than assuming a universal harm standard.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 670f9924-e8be-4008-be40-93b6ea68dad9Cited by top-tier papers4
- PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI HarmJingjing Li, Joel Mire, Eve Fleisig, Valentina Pyatkin et al.ICLR 2026 · 6 citations
- SIF: Semantically In-Distribution Fingerprints for Large Vision-Language ModelsYifei Zhao, Qian Lou, Mengxin ZhengCVPR 2026 · 2 citations
- Into the Gray Zone: Domain Contexts Can Blur LLM Safety BoundariesKi Sen Hung, Xi Yang, Chang Liu, Haoran Li et al.ACL 2026 · 1 citation
- Expectation Alignment of Language Models for Real-World User ExpectationsMiaomiao Li, Yang Wang, Bin Liang, Shudong Liu et al.ICML 2026
Builds on16
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text DataXuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel et al.UbiComp 2024 · 281 citations
- Process for Adapting Language Models to Society (PALMS) with Values-Targeted DatasetsIrene Solaiman, Christy DennisonNeurIPS 2021 · 276 citations
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
- Time Travel in LLMs: Tracing Data Contamination in Large Language ModelsShahriar Golchin, Mihai SurdeanuICLR 2024 · 165 citations
Related papers
- SafeSci: Safety Evaluation of Large Language Models in Science Domains and BeyondXiangyang Zhu, Yuan Tian, Qi Jia, Kaiwei Zhang et al.ICML 2026 · 1 citation
- Unintended Harms of Value-Aligned LLMs: Psychological and Empirical InsightsSooyung Choi, Jaehyeok Lee, Xiaoyuan Yi, Jing Yao et al.ACL 2025
- Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language ModelsYingshui Tan, Boren Zheng, Baihui Zheng, Kerui Cao et al.ACL 2025 · 7 citations
- PersonalLLM: Tailoring LLMs to Individual PreferencesThomas P. Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li et al.ICLR 2025
- OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always!Jingdi Lei, Varun Gumma, Rishabh Bhardwaj, Seok Min Lim et al.ICLR 2026 · 5 citations
