PLASH: Provably Linear-Time Attention with Selective Higher-Order Feature Sketching
Yuwen Huang, Xiang Pan
摘要
Standard softmax attention scales quadratically with sequence length, which makes long-context training and inference expensive. We introduce PLASH, an attention block whose cost grows linearly in the number of keys. PLASH compresses the original keys and values into learned prototypes, where is much smaller than the number of keys. The compressed prototypes are then enriched with randomized polynomial features that recover inter-token information lost to compression. The output is computed by exact scaled dot-product softmax attention from each query to the enriched prototypes, so PLASH preserves the standard attention interface.The construction applies to self- and cross-attention. We prove sketch-error bounds for the enrichment step, a per-input certificate that upper-bounds the deviation from standard softmax attention on each forward pass, and a runtime bound linear in the number of queries and keys. Experiments on long-context language modeling (Qwen3-4B on PG-19) and time-series forecasting (ETT, ECL, Weather) show competitive accuracy and favorable scaling against efficient-attention baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu 等ICML 2023 · 被引用 481 次
相关 Paper
- RACE Attention: A Strictly Linear-Time Attention for Long-Sequence TrainingSahil Joshi, Agniva Chowdhury, Amar Kanakamedala, Ekam Singh 等ICLR 2026 · 被引用 2 次
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token SelectionDongwon Jo, Beomseok Kang, Jiwon Song, jae-joon kimICML 2026 · 被引用 1 次
- PolySketchFormer: Fast Transformers via Sketching Polynomial KernelsPraneeth Kacham, Vahab Mirrokni, Peilin ZhongICML 2024 · 被引用 27 次
- A Unified Sparse Attention via Multi-Granularity CompressionSiran Liu, Zheng Cao, Yongchao HeICML 2026
- MVA: Linear Attention with High-order Query-Keys Integration and Multi-level Vocabulary DecompositionNing Wang, Zekun Li, Tongxin Bai, Man Yao 等ICML 2025
