Tracing the Persona Circuit: How Large Language Models Encode and Express Character Traits
Guanzheng Qin, Chenghao Sun, Zhining Xie, Xinmei Tian
摘要
Large Language Models (LLMs) demonstrate remarkable potential in role-playing tasks but frequently suffer from personality decay—termed "Out-of-Character" (OOC) behavior—during prolonged interactions. While heuristic strategies exist to align model behaviors, the internal computational dynamics driving personality expression remain opaque. A fundamental barrier to decoding these mechanisms is a metric gap: while standard causal attribution paradigms target atomic, single-token outcomes, personality manifests as a holistic, multi-token behavioral tendency. We bridge this gap via the Latent Persona Vector, a differentiable proxy enabling the first fine-grained causal tracing of personality circuits. This metric reveals a structured "Preparation-Establishment-Expression" dynamic and identifies a mechanistic contributor to OOC behavior: competition between persona-specific signals and an assistant-like default direction during the critical "Establishment" phase. Guided by this diagnosis, we propose surgically recalibrating the signal magnitude in fewer than of attention heads. This targeted intervention effectively strengthens the persona signal, significantly restoring character consistency while preserving general reasoning capabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
相关 Paper
- The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMsPengrui Han, Rafal Kocielnik, Peiyang Song, Ramit Debnath 等ICML 2026 · 被引用 33 次
- PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector AlgebraXiachong Feng, Liang Zhao, Weihong Zhong, Yichong Huang 等ICLR 2026 · 被引用 11 次
- Inertia in Moral and Value Judgments of Large Language ModelsBruce W. Lee, Yeongheon Lee, Hyunsoo ChoACL 2026 · 被引用 5 次
- Beyond Static Persona Consistency: Dynamic Persona Coherence in LLM Role-PlayingYirui Qi, Xiaoming Zhang, Ruilin Zeng, Mengyao Liu 等ACL 2026
- Can LLM Agents Maintain a Persona in Discourse?Pranav Bhandari, Nicolas Fay, Michael J. Wise, Amitava Datta 等EMNLP 2025
