Codebook Features: Sparse and Discrete Interpretability for Neural Networks
Alex Tamkin, Mohammad Taufeeque, Noah D. Goodman
摘要
Understanding neural networks is challenging in part because of the dense, continuous nature of their hidden states. We explore whether we can train neural networks to have hidden states that are sparse, discrete, and more interpretable by quantizing their continuous features into what we call codebook features. Codebook features are produced by finetuning neural networks with vector quantization bottlenecks at each layer, producing a network whose hidden features are the sum of a small number of discrete vector codes chosen from a larger codebook. Surprisingly, we find that neural networks can operate under this extreme bottleneck with only modest degradation in performance. This sparse, discrete bottleneck also provides an intuitive way of controlling neural network behavior: first, find codes that activate when the desired behavior is present, then activate those same codes during generation to elicit that behavior. We validate our approach by training codebook Transformers on several different datasets. First, we explore a finite state machine dataset with far more hidden states than neurons. In this setting, our approach overcomes the superposition problem by assigning states to distinct codes, and we find that we can make the neural network behave as if it is in a different state by activating the code for that state. Second, we train Transformer language models with up to 410M parameters on two natural language datasets. We identify codes in these models representing diverse, disentangled concepts (ranging from negative emotions to months of the year) and find that we can guide the model to generate different topics by activating the appropriate codes during inference. Overall, codebook features appear to be a promising unit of analysis and control for neural networks and interpretability. Our codebase and models are open-sourced at https://github.com/taufeeque9/codebook-features 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Identifying Functionally Important Features with End-to-End Sparse Dictionary LearningDan Braun, Jordan Taylor, Nicholas Goldowsky-Dill, Lee SharkeyNeurIPS 2024 · 被引用 81 次
- Improving Sparse Decomposition of Language Model Activations with Gated Sparse AutoencodersSenthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum 等NeurIPS 2024 · 被引用 49 次
- Obfuscated Activations Bypass LLM Latent-Space DefensesLuke Bailey, Alex Serrano, Abhay Sheshadri, Mikhail Seleznyov 等ICLR 2026 · 被引用 28 次
- NeuroStrike: Neuron-Level Attacks on Aligned LLMsLichao Wu, Sasha Behrouzi, Mohamadreza Rostami, Maximilian Thang 等NDSS 2026 · 被引用 21 次
- InversionView: A General-Purpose Method for Reading Information from Neural ActivationsXinting Huang, Madhur Panwar, Navin Goyal, Michael HahnNeurIPS 2024 · 被引用 10 次
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
相关 Paper
- Beyond Concept Bottleneck Models: How to Make Black Boxes Intervenable?Sonia Laguna, Ricards Marcinkevics, Moritz Vandenhirtz, Julia E. VogtNeurIPS 2024 · 被引用 39 次
- Explanation Bottleneck ModelsShin'ya Yamaguchi, Kosuke NishidaAAAI 2025 · 被引用 4 次
- Automatically Interpreting Millions of Features in Large Language ModelsGonçalo Paulo, Alex Mallen, Caden Juang, Nora BelroseICML 2025
- Internal Planning in Language Models: Characterizing Horizon and Branch AwarenessMuhammed Ustaomeroglu, Baris Askin, Gauri Joshi, Carlee Joe-Wong 等ICLR 2026
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsSamuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov 等ICLR 2025
