The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
Pengrui Han, Rafal Kocielnik, Peiyang Song, Ramit Debnath, Dean Mobbs, Anima Anandkumar, R. Michael Alvarez
Abstract
Personality traits have long been studied as predictors of human behavior. Recent advances in Large Language Models (LLMs) suggest similar patterns may emerge in artificial systems, with advanced LLMs displaying consistent behavioral tendencies resembling human traits like agreeableness and self-regulation. Understanding these patterns is crucial, yet prior work primarily relied on simplified self-reports and heuristic prompting, with little behavioral validation. In this study, we systematically characterize LLM personality across three dimensions: (1) the dynamic emergence and evolution of trait profiles throughout training stages; (2) the predictive validity of self-reported traits in behavioral tasks; and (3) the impact of targeted interventions, such as persona injection, on both self-reports and behavior. Our findings reveal that instructional alignment (e.g., RLHF, instruction tuning) significantly stabilizes trait expression and strengthens trait correlations in ways that mirror human data. However, these self-reported traits do not reliably predict behavior, and observed associations often diverge from human patterns. While persona injection successfully steers self-reports in the intended direction, it exerts little or inconsistent effect on actual behavior. By distinguishing surface-level trait expression from behavioral consistency, our findings challenge assumptions about LLM personality and underscore the need for deeper evaluation in alignment and interpretability. We make public all code and source data at https://github.com/ psychology-of-AI/Personality-Illusion for full transparency and reproducibility, to benefit future works in this direction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f05da5c-c238-4771-b49a-d4ea5f1128d9Cited by top-tier papers6
- PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector AlgebraXiachong Feng, Liang Zhao, Weihong Zhong, Yichong Huang et al.ICLR 2026 · 11 citations
- CoBRA: Programming Cognitive Bias in Social Agents Using Classic Social Science ExperimentsXuan Liu, HaoYang Shang, Haojian JinCHI 2026 · 2 citations
- Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented TasksHasibur Rahman, Smit DesaiCHI 2026 · 1 citation
- Persona-E²: A Human-Grounded Dataset for Personality-Shaped Emotional Responses to Textual EventsYuqin Yang, Haowu Zhou, Haoran Tu, Zhiwen Hui et al.ACL 2026
- HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive PatternsXintao Wang, Jian Yang, Weiyuan Li, Rui Xie et al.ACL 2026
Builds on19
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
Related papers
- Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language ModelsArya Shah, Deepali Mishra, Chaklam SilpasuwanchaiACL 2026 · 1 citation
- Personality Vector: Modulating Personality of Large Language Models by Model MergingSeungjong Sun, Seo Yeon Baek, Jang Hyun KimEMNLP 2025 · 5 citations
- Personality-Guided Code Generation Using Large Language ModelsYaoqi Guo, Zhenpeng Chen, Jie M. Zhang, Yang Liu et al.ACL 2025
- Chameleon LLMs: User Personas Influence Chatbot Personality ShiftsJane Xing, Tianyi Niu, Shashank SrivastavaEMNLP 2025
- Consistently Simulating Human Personas with Multi-Turn Reinforcement LearningMarwa Abdulhai, Ryan Cheng, Donovan Clay, Tim Althoff et al.NeurIPS 2025 · 51 citations
