Linear Representations of Political Perspective Emerge in Large Language Models
Junsol Kim, James Evans, Aaron Schein
摘要
Large language models (LLMs) have demonstrated the ability to generate text that realistically reflects a range of different subjective human perspectives. This paper studies how LLMs are seemingly able to reflect more liberal versus more conservative viewpoints among other political perspectives in American politics. We show that LLMs possess linear representations of political perspectives within activation space, wherein more similar perspectives are represented closer together. To do so, we probe the attention heads across the layers of three open transformerbased LLMs (Llama-2-7b-chat, Mistral-7b-instruct, Vicuna-7b). We first prompt models to generate text from the perspectives of different U.S. lawmakers. We then identify sets of attention heads whose activations linearly predict those lawmakers' DW-NOMINATE scores, a widely-used and validated measure of political ideology. We find that highly predictive heads are primarily located in the middle layers, often speculated to encode high-level concepts and tasks. Using probes only trained to predict lawmakers' ideology, we then show that the same probes can predict measures of news outlets' slant from the activations of models prompted to simulate text from those news outlets. These linear probes allow us to visualize, interpret, and monitor ideological stances implicitly adopted by an LLM as it generates open-ended responses. Finally, we demonstrate that by applying linear interventions to these attention heads, we can steer the model outputs toward a more liberal or conservative stance. Overall, our research suggests that LLMs possess a high-level linear representation of American political ideology and that by leveraging recent advances in mechanistic interpretability, we can identify, monitor, and steer the subjective perspective underlying generated text. PRELIMINARIES In this section, we define notation and provide relevant background on the architecture of transformerbased LLMs and probing methodology for discovering representations of concepts in LLMs. (5) 1 This representation is a simplification that elides details about layer normalization among other steps that are not important for the present study. However, we note that x ℓ,h will be taken before any layer normalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and RectificationStefan Krsteski, Giuseppe Russo, Serina Chang, Robert West 等ACL 2026 · 被引用 10 次
- Reinforcement Learning Fine-Tuning Enhances Activation Intensity and Diversity in the Internal Circuitry of LLMsHonglin Zhang, Qianyue Hao, Fengli Xu, Yong LiICLR 2026 · 被引用 9 次
- Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow PreferencesJoshua Ashkinaze, Hua Shen, Sai Avula, Eric Gilbert 等NeurIPS 2025 · 被引用 7 次
- DISCO: Disentangled Communication Steering for Large Language ModelsMax Torop, Aria Masoomi, Masih Eskandar, Jennifer G. DyNeurIPS 2025 · 被引用 5 次
- PoliCon: Evaluating LLMs on Achieving Diverse Political Consensus ObjectivesZhaowei Zhang, Xiaobo Wang, Minghua Yi, Mengmeng Wang 等ICLR 2026 · 被引用 4 次
它引用的顶会 Paper14
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud 等ICLR 2024 · 被引用 762 次
相关 Paper
- Hidden Persuaders: LLMs' Political Leaning and Their Influence on VotersYujin Potter, Shiyang Lai, Junsol Kim, James Evans 等EMNLP 2024 · 被引用 14 次
- Reading Between the Tokens: Improving Preference Predictions through Mechanistic ForecastingSarah Ball, Simeon Allmendinger, Frauke Kreuter, Niklas KühlICML 2026
- SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language ModelsZirui He, Mingyu Jin, Bo Shen, Ali Payani 等EMNLP 2025 · 被引用 1 次
- Examining Alignment of Large Language Models through Representative Heuristics: the case of political stereotypesSullam Jeoung, Yubin Ge, Haohan Wang, Jana DiesnerICLR 2025
- Inertia in Moral and Value Judgments of Large Language ModelsBruce W. Lee, Yeongheon Lee, Hyunsoo ChoACL 2026 · 被引用 5 次
