Focus Your Attention (with Adaptive IIR Filters)
Shahar Lutati, Itamar Zimerman, Lior Wolf
摘要
We present a new layer in which dynamic (i.e., input-dependent) Infinite Impulse Response (IIR) filters of order two are used to process the input sequence prior to applying conventional attention. The input is split into chunks, and the coefficients of these filters are determined based on previous chunks to maintain causality. Despite their relatively low order, the causal adaptive filters are shown to focus attention on the relevant sequence elements. The new layer is grounded in control theory, and is shown to generalize diagonal state-space layers. The layer performs on-par with state-of-the-art networks, with a fraction of their parameters and with time complexity that is sub-quadratic with input size. The obtained layer is favorable to layers such as Heyna, GPT2, and Mega, both with respect to the number of parameters and the obtained level of performance on multiple long-range sequence problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Zoology: Measuring and Improving Recall in Efficient Language ModelsSimran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson 等ICLR 2024 · 被引用 140 次
- Transformer-VQ: Linear-Time Transformers via Vector QuantizationLucas D. LingleICLR 2024 · 被引用 30 次
- A 2-Dimensional State Space Layer for Spatial Inductive BiasEthan Baron, Itamar Zimerman, Lior WolfICLR 2024 · 被引用 18 次
- Viewing Transformers Through the Lens of Long Convolutions LayersItamar Zimerman, Lior WolfICML 2024 · 被引用 4 次
- Graph Signal Processing Meets Mamba2: Adaptive Filter Bank via Delta ModulationYehjin Shin, Seojin Kim, Noseong ParkICLR 2026
它引用的顶会 Paper20
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang 等ICML 2022 · 被引用 2,912 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- On the Parameterization and Initialization of Diagonal State Space ModelsAlbert Gu, Karan Goel, Ankit Gupta, Christopher RéNeurIPS 2022 · 被引用 690 次
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu 等ICML 2023 · 被引用 481 次
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando 等ICML 2023 · 被引用 474 次
相关 Paper
- Trading Complexity for Expressivity Through Structured Generalized Linear Token MixingErwan Fagnou, Paul Caillon, Blaise Delattre, Alexandre AllauzenICML 2026 · 被引用 1 次
- Laughing Hyena Distillery: Extracting Compact Recurrences From ConvolutionsStefano Massaroli, Michael Poli, Daniel Y. Fu, Hermann Kumbong 等NeurIPS 2023 · 被引用 31 次
- Short-Long Convolutions Help Hardware-Efficient Linear Attention to Focus on Long SequencesZicheng Liu, Siyuan Li, Li Wang, Zedong Wang 等ICML 2024 · 被引用 11 次
- H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for SequencesZhenhai Zhu, Radu SoricutACL 2021
- Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context LengthXuezhe Ma, Xiaomeng Yang, Wenhan Xiong, Beidi Chen 等NeurIPS 2024 · 被引用 63 次
