Bivariate Beta-LSTM
Kyungwoo Song, JoonHo Jang, Seungjae Shin, Il-Chul Moon
Abstract
Long Short-Term Memory (LSTM) infers the long term dependency through a cell state maintained by the input and the forget gate structures, which models a gate output as a value in [0,1] through a sigmoid function. However, due to the graduality of the sigmoid function, the sigmoid gate is not flexible in representing multi-modality or skewness. Besides, the previous models lack modeling on the correlation between the gates, which would be a new method to adopt inductive bias for a relationship between previous and current input. This paper proposes a new gate structure with the bivariate Beta distribution. The proposed gate structure enables probabilistic modeling on the gates within the LSTM cell so that the modelers can customize the cell state flow with priors and distributions. Moreover, we theoretically show the higher upper bound of the gradient compared to the sigmoid function, and we empirically observed that the bivariate Beta distribution gate structure provides higher gradient values in training. We demonstrate the effectiveness of the bivariate Beta gate structure on the sentence classification, image classification, polyphonic music modeling, and image caption generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Multi-timescale Representation Learning in LSTM Language ModelsShivangi Mahto, Vy Ai Vo, Javier S. Turek, Alexander HuthICLR 2021 · 33 citations
- Fast Saturating Gate for Learning Long Time Scales with Recurrent Neural NetworksKentaro Ohno, Sekitoshi Kanai, Yasutoshi IdaAAAI 2023 · 1 citation
- Multi-Perspective LSTM for Joint Visual Representation LearningAlireza Sepas-Moghaddam, Fernando Pereira, Paulo Lobato Correia, Ali EtemadCVPR 2021
- Structured Sparsification of Gated Recurrent Neural NetworksEkaterina Lobacheva, Nadezhda Chirkova, Alexander Markovich, Dmitry P. VetrovAAAI 2020 · 3 citations
- Do RNN and LSTM have Long Memory?Jingyu Zhao, Feiqing Huang, Jia Lv, Yanjie Duan et al.ICML 2020 · 184 citations
