Lune

ICML2026顶会

The Cylindrical Representation Hypothesis for Language Model Steering

Lang Gao, Jinghui Zhang, Wei Liu, Fengxian Ji, Chenxi Wang, Zirui Song, Akash Ghosh, Youssef Mohamed, Preslav Nakov, Xiuying Chen

2026年份

摘要

Steering is widely used for controlling large language models, yet its effects are often unstable and difficult to predict. Existing theoretical accounts are largely based on the Linear Representation Hypothesis (LRH), which assumes that concepts can be orthogonalized for lossless control. However, this assumption rarely holds in practice and cannot explain the variability of steering outcomes. We propose the Cylindrical Representation Hypothesis (CRH), a geometric extension of LRH that relaxes the orthogonality assumption while preserving linear concept representations. We show that overlapping concept contributions naturally induce a sample-specific cylindrical structure consisting of a central axis, a normal plane, and sensitive sectors. The central axis captures the primary semantic transition associated with a target concept, while the normal plane governs steering sensitivity. Within this plane, some sectors facilitate concept activation, while others suppress or delay it. CRH reveals an asymmetry in steering predictability: the normal plane can be inferred from difference vectors, but the sensitive sectors cannot, introducing an intrinsic source of uncertainty. This explains why steering outcomes vary across samples even when intervention directions are well aligned. Experiments spanning 100 concepts, multiple models, and diverse steering methods provide consistent evidence for the predicted cylindrical structure, suggesting that steering variability arises from representation geometry rather than imperfect steering vectors. Our code is available at: https://github.com/mbzuai-nlp/CRH.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper14

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖