Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
Amin Banayeeanzade, Ala N. Tak, Fatemeh Bahrani, Anahita Bolourani, Leonardo Blas, Emilio Ferrara, Jonathan Gratch, Sai Praneeth Karimireddy
摘要
The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interactions in socially interactive settings. We introduce PsySET, a Psychologicallyinformed benchmark to evaluate LLM Steering Effectiveness and Trustworthiness across the emotion and personality domains. Our study spans four models from different LLM families paired with various steering strategies, including prompting, fine-tuning, and representation engineering. Our results indicate that prompting is consistently effective but limited in intensity control, whereas vector injections achieve finer controllability while slightly reducing output quality. Moreover, we explore the trustworthiness of steered LLMs by assessing safety, truthfulness, fairness, and ethics, highlighting potential side effects and behavioral shifts. Notably, we observe idiosyncratic effects; for instance, even a positive emotion like joy can degrade robustness to adversarial factuality, lower privacy awareness, and increase preferential bias. Meanwhile, anger predictably elevates toxicity yet strengthens leakage resistance. Our framework establishes the first holistic evaluation of emotion and personality steering, offering insights into its interpretability and reliability for socially interactive applications. Warning: Appendix contains LLM-generated content that some may find offensive. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- Evaluating and Inducing Personality in Pre-trained Language ModelsGuangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han 等NeurIPS 2023 · 被引用 192 次
相关 Paper
- Exploring the Impact of Personality Traits on LLM Bias and ToxicityShuo Wang, Renhao Li, Xi Chen, Yulin Yuan 等EMNLP 2025 · 被引用 2 次
- How Controllable Are Large Language Models? A Unified Evaluation across Behavioral GranularitiesZiwen Xu, Kewei Xu, Haoming Xu, Haiwen Hong 等ACL 2026
- Chameleon LLMs: User Personas Influence Chatbot Personality ShiftsJane Xing, Tianyi Niu, Shashank SrivastavaEMNLP 2025
- On the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMsJen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam 等ICLR 2024 · 被引用 85 次
- The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMsPengrui Han, Rafal Kocielnik, Peiyang Song, Ramit Debnath 等ICML 2026 · 被引用 33 次
