Kronecker Attention Networks
Hongyang Gao, Zhengyang Wang, Shuiwang Ji
摘要
Attention operators have been applied on both 1-D data like texts and higher-order data such as images and videos. Use of attention operators on high-order data requires flattening of the spatial or spatial-temporal dimensions into a vector, which is assumed to follow a multivariate normal distribution. This not only incurs excessive requirements on computational resources, but also fails to preserve structures in data. In this work, we propose to avoid flattening by assuming the data follow matrix-variate normal distributions. Based on this new view, we develop Kronecker attention operators (KAOs) that operate on high-order tensor data directly. More importantly, the proposed KAOs lead to dramatic reductions in computational resources. Experimental results show that our methods reduce the amount of required computational resources by a factor of hundreds, with larger factors for higher-dimensional and higher-order data. Results also show that networks with KAOs outperform models without attention, while achieving competitive performance as those with original attention operators. CCS CONCEPTS • Computing methodologies → Artificial intelligence; Machine learning algorithms; Neural networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SeaFormer: Squeeze-enhanced Axial Transformer for Mobile Semantic SegmentationQiang Wan, Zilong Huang, Jiachen Lu, Gang Yu 等ICLR 2023 · 被引用 82 次
- LaViP: Language-Grounded Visual PromptingNilakshan Kunananthaseelan, Jing Zhang, Mehrtash HarandiAAAI 2024 · 被引用 6 次
- TensorLens: End-to-End Transformer Analysis via High-Order Attention TensorsIdo Andrew Atad, Itamar Zimerman, Shahar Katz, Lior WolfACL 2026 · 被引用 1 次
- Kronecker Generative Networks: A General Neural Architecture for Parameter-Efficient Learning Across Classification TasksYang Yang, Zhengmin Kong, Yuan Liu, Tao Huang 等ICML 2026
它引用的顶会 Paper1
相关 Paper
- COSA:Co-Operative Systolic Arrays for Multi-head Attention Mechanism in Neural Network using Hybrid Data Reuse and Fusion MethodologiesZhican Wang, Gang Wang, Honglan Jiang, Ningyi Xu 等DAC 2023 · 被引用 14 次
- Interpretable Dynamic Network Modeling of Tensor Time Series via Kronecker Time-Varying Graphical LassoShingo Higashiguchi, Koki Kawabata, Yasuko Matsubara, Yasushi SakuraiWWW 2026
- Uncovering Nested Data Parallelism and Data Reuse in DNN Computation with FractalTensorSiran Liu, Chengxiang Qi, Ying Cao, Chao Yang 等SOSP 2024 · 被引用 1 次
- Towards a General Attention Framework on Gyrovector Spaces for Matrix ManifoldsRui Wang, Chen Hu, Xiaoning Song, Xiaojun Wu 等NeurIPS 2025 · 被引用 5 次
- Convolutional Neural Network Compression through Generalized Kronecker Product DecompositionMarawan Gamal Abdel Hameed, Marzieh S. Tahaei, Ali Mosleh, Vahid Partovi NiaAAAI 2022 · 被引用 33 次
