Split Personality Training: Revealing Latent Knowledge Through Alternate Personalities
Florian Dietz, William Wale, Oscar Gilg, Robert McCarthy, Felix Michalak, Gustavo Danon, Miguelito De Guzman, Dietrich Klakow
摘要
Detecting misalignment in large language models is challenging because models may learn to conceal misbehavior during training. Standard auditing techniques fall short: black-box methods often cannot distinguish misaligned outputs from benign ones, and mechanistic interpretability does not scale with model capabilities. We introduce Split Personality Training (SPT), which fine-tunes a second ``honest persona'' into LoRA parameters that remain inactive during normal operation. After the main model responds, we activate the LoRA adapter and insert a trigger string, enabling the honest persona to review the response while accessing the main model's latent states. We test our method on the Anthropic Auditing Game Model Organism, a benchmark where Llama-3.3-70B is trained to exploit reward hacks while concealing this behavior. SPT achieves 96% overall accuracy, whereas Anthropic reports near 0% accuracy. The honest persona reveals latent knowledge inaccessible to external observers, such as the fictional biases the compromised model was trained on.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen 等CCS 2024 · 被引用 132 次
- SelfIE: Self-Interpretation of Large Language Model EmbeddingsHaozhe Chen, Carl Vondrick, Chengzhi MaoICML 2024 · 被引用 58 次
- Discovering Latent Knowledge in Language Models Without SupervisionCollin Burns, Haotian Ye, Dan Klein, Jacob SteinhardtICLR 2023 · 被引用 45 次
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation ExplainersAdam Karvonen, James Chua, Clément Dumas, Kit Fraser-Taliente 等ICML 2026 · 被引用 42 次
相关 Paper
- Introspection Adapters: Training LLMs to Report Their Learned BehaviorsKeshav Shenoy, Li Yang, Abhay Sheshadri, Jack Lindsey 等ICML 2026 · 被引用 5 次
- Tracing the Persona Circuit: How Large Language Models Encode and Express Character TraitsGuanzheng Qin, Chenghao Sun, Zhining Xie, Xinmei TianICML 2026
- Your Language Model Secretly Contains Personality SubnetworksRuimeng Ye, Zihan Wang, Zinan Ling, Yang Xiao 等ICLR 2026 · 被引用 4 次
- The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMsPengrui Han, Rafal Kocielnik, Peiyang Song, Ramit Debnath 等ICML 2026 · 被引用 33 次
- PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector AlgebraXiachong Feng, Liang Zhao, Weihong Zhong, Yichong Huang 等ICLR 2026 · 被引用 11 次
