The Linear Representation Hypothesis and the Geometry of Large Language Models
Kiho Park, Yo Joong Choe, Victor Veitch
摘要
Informally, the "linear representation hypothesis" is the idea that high-level concepts are represented linearly as directions in some representation space. In this paper, we address two closely related questions: What does "linear representation" actually mean? And, how do we make sense of geometric notions (e.g., cosine similarity and projection) in the representation space? To answer these, we use the language of counterfactuals to give two formalizations of linear representation, one in the output (word) representation space, and one in the input (context) space. We then prove that these connect to linear probing and model steering, respectively. To make sense of geometric notions, we use the formalization to identify a particular (non-Euclidean) inner product that respects language structure in a sense we make precise. Using this causal inner product, we show how to unify all notions of linear representation. In particular, this allows the construction of probes and steering vectors using counterfactual pairs. Experiments with LLaMA-2 demonstrate the existence of linear representations of concepts, the connection to interpretation and control, and the fundamental role of the choice of inner product. Code is available at github.com/KihoPark/linear rep geometry.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper264
- Implicit In-context LearningZhuowei Li, Zihao Xu, Ligong Han, Yunhe Gao 等ICLR 2025 · 被引用 1,989 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- ReFT: Representation Finetuning for Language ModelsZhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger 等NeurIPS 2024 · 被引用 233 次
- A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and ToxicityAndrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg 等ICML 2024 · 被引用 177 次
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersDavid Chanin, James Wilken-Smith, Tomás Dulka, Hardik Bhatnagar 等NeurIPS 2025 · 被引用 168 次
它引用的顶会 Paper13
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 被引用 303 次
- Contrastive Learning Inverts the Data Generating ProcessRoland S. Zimmermann, Yash Sharma, Steffen Schneider, Matthias Bethge 等ICML 2021 · 被引用 264 次
- Function Vectors in Large Language ModelsEric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller 等ICLR 2024 · 被引用 229 次
相关 Paper
- The Geometry of Categorical and Hierarchical Concepts in Large Language ModelsKiho Park, Yo Joong Choe, Yibo Jiang, Victor VeitchICLR 2025 · 被引用 3 次
- On the Origins of Linear Representations in Large Language ModelsYibo Jiang, Goutham Rajendran, Pradeep Kumar Ravikumar, Bryon Aragam 等ICML 2024 · 被引用 68 次
- The Lattice Representation Hypothesis of Large Language ModelsBo XiongICLR 2026 · 被引用 3 次
- The Cylindrical Representation Hypothesis for Language Model SteeringLang Gao, Jinghui Zhang, Wei Liu, Fengxian Ji 等ICML 2026
- The Information Geometry of Softmax: Probing and SteeringKiho Park, Todd Nief, Yo Joong Choe, Victor VeitchICML 2026 · 被引用 6 次
