Bilinear MLPs enable weight-based mechanistic interpretability
Michael T. Pearce, Thomas Dooms, Alice Rigg, José Oramas, Lee Sharkey
摘要
A mechanistic understanding of how MLPs do computation in deep neural networks remains elusive. Current interpretability work can extract features from hidden activations over an input dataset but generally cannot explain how MLP weights construct features. One challenge is that element-wise nonlinearities introduce higher-order interactions and make it difficult to trace computations through the MLP layer. In this paper, we analyze bilinear MLPs, a type of Gated Linear Unit (GLU) without any element-wise nonlinearity that nevertheless achieves competitive performance. Bilinear MLPs can be fully expressed in terms of linear operations using a third-order tensor, allowing flexible analysis of the weights. Analyzing the spectra of bilinear MLP weights using eigendecomposition reveals interpretable low-rank structure across toy tasks, image classification, and language modeling. We use this understanding to craft adversarial examples, uncover overfitting, and identify small language model circuits directly from the weights alone. Our results demonstrate that bilinear layers serve as an interpretable drop-in replacement for current activation functions and that weightbased interpretability is viable for understanding deep-learning models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Beyond Linear Probes: Dynamic Safety Monitoring for Language ModelsJames Oldfield, Philip Torr, Ioannis Patras, Adel Bibi 等ICLR 2026 · 被引用 16 次
- Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of DecodersJames Oldfield, Shawn Im, Sharon Li, Mihalis A. Nicolaou 等NeurIPS 2025 · 被引用 7 次
- Volume Transmission Implements Context Factorization to Target Online Credit Assignment and Enable Compositional GeneralizationMatthew S. Bull, Po-Chen Kuo, Andrew L. Smith, Michael A. BuiceNeurIPS 2025 · 被引用 4 次
- Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and ArchitecturesYutong Gao, Qinglin Meng, Yuan Zhou, Liangming PanACL 2026 · 被引用 3 次
- PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial DecodingPanagiotis Koromilas, Andreas Demou, James Oldfield, Yannis Panagakis 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper5
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Emergence of Sparse Representations from NoiseTrenton Bricken, Rylan Schaeffer, Bruno A. Olshausen, Gabriel KreimanICML 2023 · 被引用 15 次
- Scaling and evaluating sparse autoencodersLeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh 等ICLR 2025 · 被引用 10 次
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsSamuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov 等ICLR 2025
相关 Paper
- Understanding MLP-Mixer as a wide and sparse MLPTomohiro Hayase, Ryo KarakidaICML 2024 · 被引用 9 次
- Beyond Components: Singular Vector-Based Interpretability of Transformer CircuitsAreeb Ahmad, Abhinav Joshi, Ashutosh ModiNeurIPS 2025 · 被引用 9 次
- KAN: Kolmogorov-Arnold NetworksZiming Liu, Yixuan Wang, Sachin Vaidya, Fabian Ruehle 等ICLR 2025
- GmNet: Revisiting Gating Mechanisms From A Frequency ViewYifan Wang, Xu Ma, Yitian Zhang, Yizhou Wang 等ICLR 2026 · 被引用 1 次
- Deep Learning with Learnable Product-Structured ActivationsSaanjali Maharaj, Prasanth B. NairICLR 2026
