Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution
Haiyan Zhao, Heng Zhao, Bo Shen, Ali Payani, Fan Yang, Mengnan Du
摘要
Probing learned concepts in large language models (LLMs) is crucial for understanding how semantic knowledge is encoded internally. Training linear classifiers on probing tasks is a principle approach to denote the vector of a certain concept in the representation space. However, the single vector identified for a concept varies with both data and training, making it less robust and weakening its effectiveness in real-world applications. To address this challenge, we propose an approach to approximate the subspace representing a specific concept. Built on linear probing classifiers, we extend the concept vectors into Gaussian Concept Subspace (GCS). We demonstrate GCS's effectiveness through measuring its faithfulness and plausibility across multiple LLMs with different sizes and architectures. Additionally, we use representation intervention tasks to showcase its efficacy in real-world applications such as emotion steering. Experimental results indicate that GCS concept vectors have the potential to balance steering performance and maintaining the fluency in natural language generation tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Mechanism of Task-oriented Information Removal in In-context LearningHakaze Cho, Haolin Yang, Gouki Minegishi, Naoya InoueICLR 2026 · 被引用 3 次
- The Lattice Representation Hypothesis of Large Language ModelsBo XiongICLR 2026 · 被引用 3 次
- Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head AnalysisHaolin Yang, Hakaze Cho, Naoya InoueICLR 2026 · 被引用 2 次
- Cell-Based Representation of Relational Binding in Language ModelsQin Dai, Benjamin Heinzerling, Kentaro InuiACL 2026 · 被引用 1 次
- Neuron-Level Differentiation of Memorization and Generalization in Large Language ModelsKo-Wei Huang, Yi-Fu Fu, Ching-Yu Tsai, Yu-Chieh Tu 等EMNLP 2025
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
相关 Paper
- Gaussian Process Probes (GPP) for Uncertainty-Aware ProbingZi Wang, Alexander Ku, Jason Baldridge, Tom Griffiths 等NeurIPS 2023 · 被引用 17 次
- SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language ModelsZirui He, Mingyu Jin, Bo Shen, Ali Payani 等EMNLP 2025 · 被引用 1 次
- Towards Understanding Steering StrengthMagamed Taimeskhanov, Samuel Vaiter, Damien GarreauICML 2026 · 被引用 2 次
- Unsupervised Concept Vector Extraction for Bias Control in LLMsHannah Cyberey, Yangfeng Ji, David EvansEMNLP 2025 · 被引用 4 次
- Improving Preference Extraction In LLMs By Identifying Latent Knowledge Through Classifying ProbesSharan Maiya, Yinhong Liu, Ramit Debnath, Anna KorhonenACL 2025 · 被引用 4 次
