Kernel Multimodal Continuous Attention
Alexander Moreno, Zhenke Wu, Supriya Nagesh, Walter H. Dempsey, James M. Rehg
Abstract
Attention mechanisms average a data representation with respect to probability weights. Recently, [23] [24] [25] proposed continuous attention, focusing on unimodal exponential and deformed exponential family attention densities: the latter can have sparse support. [8] extended to multimodality via Gaussian mixture attention densities. In this paper, we propose using kernel exponential families [4] and our new sparse counterpart, kernel deformed exponential families. Theoretically, we show new existence results for both families, and approximation capabilities for the deformed case. Lacking closed form expressions for the context vector, we use numerical integration: we prove exponential convergence for both families. Experiments show that kernel continuous attention often outperforms unimodal continuous attention, and the sparse variant tends to highlight time series peaks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Multi-Time Attention Networks for Irregularly Sampled Time SeriesSatya Narayan Shukla, Benjamin M. MarlinICLR 2021 · 301 citations
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song et al.ICLR 2021 · 122 citations
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty et al.KDD 2021 · 66 citations
- Sparse and Continuous Attention MechanismsAndré F. T. Martins, António Farinhas, Marcos V. Treviso, Vlad Niculae et al.NeurIPS 2020 · 55 citations
- Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time SeriesSatya Narayan Shukla, Benjamin M. MarlinICLR 2022 · 29 citations
Related papers
- Non-stationary Time-aware Kernelized Attention for Temporal Event PredictionYu Ma, Zhining Liu, Chenyi Zhuang, Yize Tan et al.KDD 2022 · 4 citations
- Continuous-Time Attention for Sequential LearningJen-Tzung Chien, Yi-Hsiang ChenAAAI 2021 · 19 citations
- FourierFormer: Transformer Meets Generalized Fourier Integral TheoremTan Nguyen, Minh Pham, Tam Nguyen, Khai Nguyen et al.NeurIPS 2022 · 59 citations
- A Temporal Kernel Approach for Deep Learning with Continuous-time InformationDa Xu, Chuanwei Ruan, Evren Körpeoglu, Sushant Kumar et al.ICLR 2021 · 7 citations
- Sequential Recommendation with Relation-Aware Kernelized Self-AttentionMingi Ji, Weonyoung Joo, Kyungwoo Song, Yoon-Yeong Kim et al.AAAI 2020 · 31 citations
