Recasting Self-Attention with Holographic Reduced Representations
Mohammad Mahmudul Alam, Edward Raff, Stella Biderman, Tim Oates, James Holt
Abstract
In recent years, self-attention has become the dominant paradigm for sequence modeling in a variety of domains. However, in domains with very long sequence lengths the memory and compute costs can make using transformers infeasible. Motivated by problems in malware detection, where sequence lengths of are a roadblock to deep learning, we re-cast self-attention using the neuro-symbolic approach of Holographic Reduced Representations (HRR). In doing so we perform the same high-level strategy of the standard self-attention: a set of queries matching against a set of keys, and returning a weighted response of the values for each key. Implemented as a ``Hrrformer'' we obtain several benefits including time complexity, space complexity, and convergence in fewer epochs. Nevertheless, the Hrrformer achieves near state-of-the-art accuracy on LRA benchmarks and we are able to learn with just a single layer. Combined, these benefits make our Hrrformer the first viable Transformer for such long malware classification sequences and up to faster to train on the Long Range Arena benchmark. Code is available at https://github.com/NeuromorphicComputationResearchProgram/Hrrformer
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- MIMONets: Multiple-Input-Multiple-Output Neural Networks Exploiting Computation in SuperpositionNicolas Menet, Michael Hersche, Geethan Karunaratne, Luca Benini et al.NeurIPS 2023 · 32 citations
- A Walsh Hadamard Derived Linear Vector Symbolic ArchitectureMohammad Mahmudul Alam, Alexander Oberle, Edward Raff, Stella Biderman et al.NeurIPS 2024 · 10 citations
- Beyond Raw Bytes: Towards Large Malware Language ModelsLuke Kurlandski, Harel Berger, Yin Pan, Matthew WrightNDSS 2026 · 5 citations
- FoGE: Fock Space inspired encoding for graph promptingTakis Chytas, Rudrasis Chakraborty, Vikas SinghNeurIPS 2025
Builds on15
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab et al.NeurIPS 2021 · 1,280 citations
Related papers
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
- H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for SequencesZhenhai Zhu, Radu SoricutACL 2021
- Classifying Sequences of Extreme Length with Constant Memory Applied to Malware DetectionEdward Raff, William Fleshman, Richard Zak, Hyrum S. Anderson et al.AAAI 2021 · 70 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
- Treeformer: Dense Gradient Trees for Efficient Attention ComputationLovish Madaan, Srinadh Bhojanapalli, Himanshu Jain, Prateek JainICLR 2023
