Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation
Itamar Zimerman, Ameen Ali, Lior Wolf
摘要
Recent advances in efficient sequence modeling have led to attention-free layers, such as Mamba, RWKV, and various gated RNNs, all featuring sub-quadratic complexity in sequence length and excellent scaling properties, enabling the construction of a new type of foundation models. In this paper, we present a unified view of these models, formulating such layers as implicit causal self-attention layers. The formulation includes most of their sub-components and is not limited to a specific part of the architecture. The framework compares the underlying mechanisms on similar grounds for different layers and provides a direct means for applying explainability methods. Our experiments show that our attention matrices and attribution method outperform an alternative and a more limited formulation that was recently proposed for Mamba. For the other architectures for which our method is the first to provide such a view, our method is effective and competitive in the relevant metrics compared to the results obtained by state-of-the-art Transformer explainability methods. Our code is publicly available. https://github.com/Itamarzimm/UnifiedImplicitAttnRepr
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer ExplainabilityYarden Bakish, Itamar Zimerman, Hila Chefer, Lior WolfNeurIPS 2025 · 被引用 8 次
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video UnderstandingBoshen Xu, Zihan Xiao, Jiaze Li, Jianzhong Ju 等CVPR 2026 · 被引用 5 次
- State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video UnderstandingJiahuan Zhou, Kai Zhu, Zhenyu Cui, Zichen Liu 等NeurIPS 2025 · 被引用 2 次
- TensorLens: End-to-End Transformer Analysis via High-Order Attention TensorsIdo Andrew Atad, Itamar Zimerman, Shahar Katz, Lior WolfACL 2026 · 被引用 1 次
- RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic SegmentationKai Zhu, Zhenyu Cui, Zehua Zang, Jiahuan ZhouCVPR 2026 · 被引用 1 次
它引用的顶会 Paper32
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
相关 Paper
- Learning to (Learn at Test Time): RNNs with Expressive Hidden StatesYu Sun, Xinhao Li, Karan Dalal, Jiarui Xu 等ICML 2025
- Log-Linear AttentionHan Guo, Songlin Yang, Tarushii Goel, Eric P. Xing 等ICLR 2026 · 被引用 41 次
- MambaLRP: Explaining Selective State Space Sequence ModelsFarnoush Rezaei Jafari, Grégoire Montavon, Klaus-Robert Müller, Oliver EberleNeurIPS 2024 · 被引用 44 次
- The Hidden Attention of Mamba ModelsAmeen Ali, Itamar Zimerman, Lior WolfACL 2025
- Hydra: Bidirectional State Space Models Through Generalized Matrix MixersSukjun Hwang, Aakash Sunil Lahoti, Ratish Puduppully, Tri Dao 等NeurIPS 2024 · 被引用 54 次
