Kernel Multimodal Continuous Attention
Alexander Moreno, Zhenke Wu, Supriya Nagesh, Walter H. Dempsey, James M. Rehg
摘要
Attention mechanisms average a data representation with respect to probability weights. Recently, [23] [24] [25] proposed continuous attention, focusing on unimodal exponential and deformed exponential family attention densities: the latter can have sparse support. [8] extended to multimodality via Gaussian mixture attention densities. In this paper, we propose using kernel exponential families [4] and our new sparse counterpart, kernel deformed exponential families. Theoretically, we show new existence results for both families, and approximation capabilities for the deformed case. Lacking closed form expressions for the context vector, we use numerical integration: we prove exponential convergence for both families. Experiments show that kernel continuous attention often outperforms unimodal continuous attention, and the sparse variant tends to highlight time series peaks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Multi-Time Attention Networks for Irregularly Sampled Time SeriesSatya Narayan Shukla, Benjamin M. MarlinICLR 2021 · 被引用 301 次
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 等ICLR 2021 · 被引用 122 次
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty 等KDD 2021 · 被引用 66 次
- Sparse and Continuous Attention MechanismsAndré F. T. Martins, António Farinhas, Marcos V. Treviso, Vlad Niculae 等NeurIPS 2020 · 被引用 55 次
- Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time SeriesSatya Narayan Shukla, Benjamin M. MarlinICLR 2022 · 被引用 29 次
相关 Paper
- Non-stationary Time-aware Kernelized Attention for Temporal Event PredictionYu Ma, Zhining Liu, Chenyi Zhuang, Yize Tan 等KDD 2022 · 被引用 4 次
- Continuous-Time Attention for Sequential LearningJen-Tzung Chien, Yi-Hsiang ChenAAAI 2021 · 被引用 19 次
- FourierFormer: Transformer Meets Generalized Fourier Integral TheoremTan Nguyen, Minh Pham, Tam Nguyen, Khai Nguyen 等NeurIPS 2022 · 被引用 59 次
- A Temporal Kernel Approach for Deep Learning with Continuous-time InformationDa Xu, Chuanwei Ruan, Evren Körpeoglu, Sushant Kumar 等ICLR 2021 · 被引用 7 次
- Sequential Recommendation with Relation-Aware Kernelized Self-AttentionMingi Ji, Weonyoung Joo, Kyungwoo Song, Yoon-Yeong Kim 等AAAI 2020 · 被引用 31 次
