Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
Mor Geva, Avi Caciularu, Kevin Ro Wang, Yoav Goldberg
摘要
Transformer-based language models (LMs) are at the core of modern NLP, but their internal prediction construction process is opaque and largely not understood. In this work, we make a substantial step towards unveiling this underlying prediction process, by reverseengineering the operation of the feed-forward network (FFN) layers, one of the building blocks of transformer models. We view the token representation as a changing distribution over the vocabulary, and the output from each FFN layer as an additive update to that distribution. Then, we analyze the FFN updates in the vocabulary space, showing that each update can be decomposed to sub-updates corresponding to single FFN parameter vectors, each promoting concepts that are often human-interpretable. We then leverage these findings for controlling LM predictions, where we reduce the toxicity of GPT2 by almost 50%, and for improving computation efficiency with a simple early exit rule, saving 20% of computation on average. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper218
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni 等ICLR 2024 · 被引用 462 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani 等NeurIPS 2022 · 被引用 394 次
- Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language ModelsPeter Hase, Mohit Bansal, Been Kim, Asma GhandehariounNeurIPS 2023 · 被引用 307 次
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 被引用 233 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- DynaBERT: Dynamic BERT with Adaptive Width and DepthLu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang 等NeurIPS 2020 · 被引用 401 次
- Depth-Adaptive TransformerMaha Elbayad, Jiatao Gu, Edouard Grave, Michael AuliICLR 2020 · 被引用 264 次
- What Language Model Architecture and Pretraining Objective Works Best for Zero-Shot Generalization?Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao 等ICML 2022 · 被引用 228 次
- Transformer Feed-Forward Layers Are Key-Value MemoriesMor Geva, Roei Schuster, Jonathan Berant, Omer LevyEMNLP 2021 · 被引用 33 次
相关 Paper
- Token-wise Decomposition of Autoregressive Language Model Hidden States for Analyzing Model PredictionsByung-Doh Oh, William SchulerACL 2023 · 被引用 1 次
- Analyzing Vision Transformers for Image Classification in Class Embedding SpaceMartina G. Vilas, Timothy Schaumlöffel, Gemma RoigNeurIPS 2023 · 被引用 43 次
- Head Pursuit: Probing Attention Specialization in Multimodal TransformersLorenzo Basile, Valentino Maiorca, Diego Doimo, Francesco Locatello 等NeurIPS 2025 · 被引用 21 次
- LLM Braces: Straightening Out LLM Predictions with Relevant Sub-UpdatesYing Shen, Lifu HuangACL 2025 · 被引用 3 次
- Backward Lens: Projecting Language Model Gradients into the Vocabulary SpaceShahar Katz, Yonatan Belinkov, Mor Geva, Lior WolfEMNLP 2024 · 被引用 2 次
