Concept Heterogeneity-aware Representation Steering
Laziz Abdullaev, Noelle Y. L. Wong, Ryan Lee, Shiqi Jiang, Minh-Khoi Nguyen-Nhat, Tan Nguyen
摘要
Representation steering offers a lightweight mechanism for controlling the behavior of large language models (LLMs) by intervening on internal activations at inference time. Most existing methods rely on a single global steering direction, typically obtained via difference-in-means over contrastive datasets. This approach implicitly assumes that the target concept is homogeneously represented across the embedding space. In practice, however, LLM representations can be highly non-homogeneous, exhibiting clustered, context-dependent structure, which renders global steering directions brittle. In this work, we view representation steering through the lens of optimal transport (OT), noting that standard difference-in-means steering implicitly corresponds to the OT map between two identical distributions with differing first moments, yielding a global translation. To relax this restrictive assumption, we theoretically model source and target representations as Gaussian mixture models and formulate steering as a discrete OT problem between semantic latent clusters. From the resulting transport plan, we derive an explicit, inputdependent steering map via barycentric projection, producing a smooth, kernel-weighted combination of cluster-level shifts. We term this method Concept Heterogeneity-aware Representation Steering (CHaRS). Through numerous experimental settings, we show that CHaRS yields more effective behavioral control than global steering. The code is publicly available at https: //github.com/lazizcodes/CHaRS .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- Rates of Estimation of Optimal Transport Maps using Plug-in Estimators via Barycentric ProjectionsNabarun Deb, Promit Ghosal, Bodhisattva SenNeurIPS 2021 · 被引用 96 次
- Adaptive Activation Steering: A Tuning-Free LLM Truthfulness Improvement Method for Diverse Hallucinations CategoriesTianlong Wang, Xianfeng Jiao, Yinghao Zhu, Zhongzhi Chen 等WWW 2025 · 被引用 64 次
相关 Paper
- Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation SteeringEric Bigelow, Daniel Wurgaft, YingQiao Wang, Noah Goodman 等ICML 2026
- The Cylindrical Representation Hypothesis for Language Model SteeringLang Gao, Jinghui Zhang, Wei Liu, Fengxian Ji 等ICML 2026
- COLD-Steer: Steering Large Language Models via In-Context One-step Learning DynamicsKartik Sharma, Rakshit S. TrivediICLR 2026 · 被引用 8 次
- Towards Understanding Steering StrengthMagamed Taimeskhanov, Samuel Vaiter, Damien GarreauICML 2026 · 被引用 2 次
- Improved Representation Steering for Language ModelsZhengxuan Wu, Qinan Yu, Aryaman Arora, Christopher D. Manning 等NeurIPS 2025 · 被引用 22 次
