Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models
Ruixuan Deng, Xiaoyang Hu, Miles Gilberti, Shane Storks, Aman Taxali, Mike Angstadt, Chandra Sekhar Sripada, Joyce Chai
摘要
We identify semantically coherent, contextconsistent network components in large language models (LLMs) using coactivation of sparse autoencoder (SAE) features collected from just a handful of prompts. Focusing on concept-relation prediction tasks, we show that ablating these components for concepts (e.g., countries and words) and relations (e.g., capital city and translation language) changes model outputs in predictable ways, while amplifying these components induces counterfactual responses. Notably, composing relation and concept components yields compound counterfactual outputs. Further analysis reveals that while most concept components emerge from the very first layer, more abstract relation components are concentrated in later layers. Lastly, we show that extracted components more comprehensively capture concepts and relations than individual features while maintaining specificity. Overall, our findings suggest a modular organization of knowledge and advance methods for efficient, targeted LLM manipulation. 1 * Indicates equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn 等ICLR 2022 · 被引用 527 次
- How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language modelMichael Hanna, Ollie Liu, Alexandre VariengienNeurIPS 2023 · 被引用 251 次
相关 Paper
- Do Sparse Autoencoders Identify Reasoning Features in Language Models?George Ma, Zhongyuan Liang, Irene Y. Chen, Somayeh SojoudiICML 2026 · 被引用 9 次
- LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-EncoderYi Jing, Zijun Yao, Hongzhu Guo, Lingxu Ran 等EMNLP 2025 · 被引用 7 次
- Sparse Autoencoders Trained on the Same Data Learn Different FeaturesGonçalo Paulo, Nora BelroseICLR 2026 · 被引用 96 次
- ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language ModelsHaoxuan Li, Zhen Wen, Qiqi Jiang, Chenxiao Li 等IEEE VIS 2025 · 被引用 3 次
- Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMsXinwei Wu, Heng Liu, Xiaohu Zhao, Yuqi Ren 等AAAI 2026 · 被引用 2 次
