Evaluating Psychological Safety of Large Language Models
Xingxuan Li, Yutong Li, Lin Qiu, Shafiq Joty, Lidong Bing
Abstract
In this work, we designed unbiased prompts to systematically evaluate the psychological safety of large language models (LLMs). First, we tested five different LLMs by using two personality tests: Short Dark Triad (SD-3) and Big Five Inventory (BFI). All models scored higher than the human average on SD-3, suggesting a relatively darker personality pattern. Despite being instruction fine-tuned with safety metrics to reduce toxicity, InstructGPT, GPT-3.5, and GPT-4 still showed dark personality patterns; these models scored higher than self-supervised GPT-3 on the Machiavellianism and narcissism traits on SD-3. Then, we evaluated the LLMs in the GPT series by using well-being tests to study the impact of fine-tuning with more training data. We observed a continuous increase in the well-being scores of GPT models. Following these observations, we showed that finetuning Llama-2-chat-7B with responses from BFI using direct preference optimization could effectively reduce the psychological toxicity of the model. Based on the findings, we recommended the application of systematic and comprehensive psychological metrics to further evaluate and improve the safety of LLMs. 1 Warning: This paper contains examples with potentially harmful content. * Xingxuan Li is under the Joint Ph.D. Program between Alibaba and Nanyang Technological University. 1 We will make our code and data publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23b2e101-64d9-465b-b65e-ef431b780b62Cited by top-tier papers7
- On the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMsJen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam et al.ICLR 2024 · 85 citations
- Apathetic or Empathetic? Evaluating LLMs' Emotional Alignments with HumansJen-tse Huang, Man Ho Lam, Eric John Li, Shujie Ren et al.NeurIPS 2024 · 63 citations
- Measuring Human and AI Values Based on Generative Psychometrics with Large Language ModelsHaoran Ye, Yuhang Xie, Yuanyi Ren, Hanjun Fang et al.AAAI 2025 · 18 citations
- On the Reliability of Psychological Scales on Large Language ModelsJen-tse Huang, Wenxiang Jiao, Man Ho Lam, Eric John Li et al.EMNLP 2024 · 6 citations
- AI-exhibited Personality Traits Can Shape Human Self-concept through ConversationsJingshu Li, Tianqi Song, Nattapat Boonprakong, Zicheng Zhu et al.CHI 2026 · 3 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
Related papers
- Exploring the Impact of Personality Traits on LLM Bias and ToxicityShuo Wang, Renhao Li, Xi Chen, Yulin Yuan et al.EMNLP 2025 · 2 citations
- Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt TemplatesKaifeng Lyu, Haoyu Zhao, Xinran Gu, Dingli Yu et al.NeurIPS 2024 · 131 citations
- Unveiling the Implicit Toxicity in Large Language ModelsJiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang et al.EMNLP 2023 · 21 citations
- Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric StudyKensuke Okada, Yui Furukawa, Kyosuke BunjiACL 2026 · 1 citation
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger et al.ICLR 2024 · 373 citations
