The Information Geometry of Softmax: Probing and Steering
Kiho Park, Todd Nief, Yo Joong Choe, Victor Veitch
摘要
This paper concerns the question of how AI systems encode semantic structure into the geometric structure of their representation spaces. The motivating observation is that the natural geometry of these representation spaces should reflect the way models use representations to produce behavior. We focus on the important special case of representations that define softmax distributions. In this case, we argue that the natural geometry is information geometry. Our focus is on the role of information geometry on semantic encoding and the linear representation hypothesis. As an illustrative application, we develop dual steering, a method for robustly steering representations to exhibit a particular concept using linear probes. We prove that dual steering optimally modifies the target concept while minimizing changes to off-target concepts. Empirically, we find that dual steering enhances the controllability and stability of concept manipulation. Code is available at github.com/KihoPark/dual-steering.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski GeometryThomas Fel, Binxu Wang, Michael A. Lepori, Matthew Kowal 等ICLR 2026 · 被引用 28 次
- Attention Implements the Fisher Geometry of Exponential FamiliesBodie RubacherICML 2026
- How Controllable Are Large Language Models? A Unified Evaluation across Behavioral GranularitiesZiwen Xu, Kewei Xu, Haoming Xu, Haiwen Hong 等ACL 2026
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
相关 Paper
- The Cylindrical Representation Hypothesis for Language Model SteeringLang Gao, Jinghui Zhang, Wei Liu, Fengxian Ji 等ICML 2026
- Angular Steering: Behavior Control via Rotation in Activation SpaceMinh Hieu Vu, Tan M. NguyenNeurIPS 2025 · 被引用 53 次
- Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian DistributionHaiyan Zhao, Heng Zhao, Bo Shen, Ali Payani 等ICLR 2025 · 被引用 2 次
- Concept Heterogeneity-aware Representation SteeringLaziz Abdullaev, Noelle Y. L. Wong, Ryan Lee, Shiqi Jiang 等ICML 2026
- Towards Understanding Steering StrengthMagamed Taimeskhanov, Samuel Vaiter, Damien GarreauICML 2026 · 被引用 2 次
