PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra
Xiachong Feng, Liang Zhao, Weihong Zhong, Yichong Huang, Yuxuan Gu, Lingpeng Kong, Xiaocheng Feng, Bing Qin
摘要
Current methods for personality control in Large Language Models rely on static prompting or expensive fine-tuning, failing to capture the dynamic and compositional nature of human traits. We introduce PERSONA, a training-free framework that achieves fine-tuning level performance through direct manipulation of personality vectors in activation space. Our key insight is that personality traits appear as extractable, approximately orthogonal directions in the model's representation space that support algebraic operations. The framework operates through three stages: PERSONA-BASE extracts orthogonal trait vectors via contrastive activation analysis; PERSONA-ALGEBRA enables precise control through vector arithmetic (scalar multiplication for intensity, addition for composition, subtraction for suppression); and PERSONA-FLOW achieves context-aware adaptation by dynamically composing these vectors during inference. On PersonalityBench, our approach achieves a mean score of 9.60, nearly matching the supervised finetuning upper bound of 9.61 without any gradient updates. On our proposed PERSONA-EVOLVE benchmark for dynamic personality adaptation, we achieve up to 91% win rates across diverse model families. These results provide evidence that aspects of LLM personality are mathematically tractable, opening new directions for interpretable and efficient behavioral control 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
相关 Paper
- Personality Vector: Modulating Personality of Large Language Models by Model MergingSeungjong Sun, Seo Yeon Baek, Jang Hyun KimEMNLP 2025 · 被引用 5 次
- Controllable and Explainable Personality Sliders for LLMs at Inference TimeFlorian Hoppe, David Khachaturov, Robert Mullins, Mark Huasong MengICML 2026 · 被引用 1 次
- Your Language Model Secretly Contains Personality SubnetworksRuimeng Ye, Zihan Wang, Zinan Ling, Yang Xiao 等ICLR 2026 · 被引用 4 次
- Tracing the Persona Circuit: How Large Language Models Encode and Express Character TraitsGuanzheng Qin, Chenghao Sun, Zhining Xie, Xinmei TianICML 2026
- The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMsPengrui Han, Rafal Kocielnik, Peiyang Song, Ramit Debnath 等ICML 2026 · 被引用 33 次
