Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
Mor Geva, Avi Caciularu, Kevin Ro Wang, Yoav Goldberg
Abstract
Transformer-based language models (LMs) are at the core of modern NLP, but their internal prediction construction process is opaque and largely not understood. In this work, we make a substantial step towards unveiling this underlying prediction process, by reverseengineering the operation of the feed-forward network (FFN) layers, one of the building blocks of transformer models. We view the token representation as a changing distribution over the vocabulary, and the output from each FFN layer as an additive update to that distribution. Then, we analyze the FFN updates in the vocabulary space, showing that each update can be decomposed to sub-updates corresponding to single FFN parameter vectors, each promoting concepts that are often human-interpretable. We then leverage these findings for controlling LM predictions, where we reduce the toxicity of GPT2 by almost 50%, and for improving computation efficiency with a simple early exit rule, saving 20% of computation on average. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d054a4a3-d74c-40e6-ad91-3d9d1f54e250Cited by top-tier papers218
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni et al.ICLR 2024 · 462 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
- Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language ModelsPeter Hase, Mohit Bansal, Been Kim, Asma GhandehariounNeurIPS 2023 · 307 citations
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 233 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- DynaBERT: Dynamic BERT with Adaptive Width and DepthLu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang et al.NeurIPS 2020 · 401 citations
- Depth-Adaptive TransformerMaha Elbayad, Jiatao Gu, Edouard Grave, Michael AuliICLR 2020 · 264 citations
- What Language Model Architecture and Pretraining Objective Works Best for Zero-Shot Generalization?Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao et al.ICML 2022 · 228 citations
- Transformer Feed-Forward Layers Are Key-Value MemoriesMor Geva, Roei Schuster, Jonathan Berant, Omer LevyEMNLP 2021 · 33 citations
Related papers
- Token-wise Decomposition of Autoregressive Language Model Hidden States for Analyzing Model PredictionsByung-Doh Oh, William SchulerACL 2023 · 1 citation
- Analyzing Vision Transformers for Image Classification in Class Embedding SpaceMartina G. Vilas, Timothy Schaumlöffel, Gemma RoigNeurIPS 2023 · 43 citations
- Head Pursuit: Probing Attention Specialization in Multimodal TransformersLorenzo Basile, Valentino Maiorca, Diego Doimo, Francesco Locatello et al.NeurIPS 2025 · 21 citations
- LLM Braces: Straightening Out LLM Predictions with Relevant Sub-UpdatesYing Shen, Lifu HuangACL 2025 · 3 citations
- Backward Lens: Projecting Language Model Gradients into the Vocabulary SpaceShahar Katz, Yonatan Belinkov, Mor Geva, Lior WolfEMNLP 2024 · 2 citations
