Backpack Language Models
John Hewitt, John Thickstun, Christopher D. Manning, Percy Liang
Abstract
We present Backpacks: a new neural architecture that marries strong modeling performance with an interface for interpretability and control. Backpacks learn multiple non-contextual sense vectors for each word in a vocabulary, and represent a word in a sequence as a contextdependent, non-negative linear combination of sense vectors in this sequence. We find that, after training, sense vectors specialize, each encoding a different aspect of a word. We can interpret a sense vector by inspecting its (non-contextual, linear) projection onto the output space, and intervene on these interpretable hooks to change the model's behavior in predictable ways. We train a 170M-parameter Backpack language model on OpenWebText, matching the loss of a GPT-2 small (124Mparameter) Transformer. On lexical similarity evaluations, we find that Backpack sense vectors outperform even a 6B-parameter Transformer LM's word embeddings. Finally, we present simple algorithms that intervene on sense vectors to perform controllable text generation and debiasing. For example, we can edit the sense vocabulary to tend more towards a topic, or localize a source of gender bias to a sense vector and globally suppress that sense.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4eeefce2-0f07-4a23-8b40-f5e81bcb9480Cited by top-tier papers7
- Flipping the Dialogue: Training and Evaluating User Language ModelsTarek Naous, Philippe Laban, Wei Xu, Jennifer NevilleICLR 2026 · 56 citations
- Codebook Features: Sparse and Discrete Interpretability for Neural NetworksAlex Tamkin, Mohammad Taufeeque, Noah D. GoodmanICML 2024 · 42 citations
- PaCE: Parsimonious Concept Engineering for Large Language ModelsJinqi Luo, Tianjiao Ding, Kwan Ho Ryan Chan, Darshan Thaker et al.NeurIPS 2024 · 20 citations
- Word Embeddings Are Steers for Language ModelsChi Han, Jialiang Xu, Manling Li, Yi Fung et al.ACL 2024 · 8 citations
- Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and ArchitecturesYutong Gao, Qinglin Meng, Yuan Zhou, Liangming PanACL 2026 · 3 citations
Builds on13
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
Related papers
- VAST: The Valence-Assessing Semantics Test for Contextualizing Language ModelsRobert Wolfe, Aylin CaliskanAAAI 2022 · 17 citations
- How Much Do Encoder Models Know About Word Senses?Simone Teglia, Simone Tedeschi, Roberto NavigliACL 2025 · 1 citation
- Adaptive Probabilistic Word EmbeddingShuangyin Li, Yu Zhang, Rong Pan, Kaixiang MoWWW 2020 · 10 citations
- SenseBERT: Driving Some Sense into BERTYoav Levine, Barak Lenz, Or Dagan, Ori Ram et al.ACL 2020 · 27 citations
- Automatically Generated Definitions and their utility for Modeling Word MeaningFrancesco Periti, David Alfter, Nina TahmasebiEMNLP 2024 · 2 citations
