Multi-timescale Representation Learning in LSTM Language Models
Shivangi Mahto, Vy Ai Vo, Javier S. Turek, Alexander Huth
Abstract
Language models must capture statistical dependencies between words at timescales ranging from very short to very long. Earlier work has demonstrated that dependencies in natural language tend to decay with distance between words according to a power law. However, it is unclear how this knowledge can be used for analyzing or designing neural network language models. In this work, we derived a theory for how the memory gating mechanism in long short-term memory (LSTM) language models can capture power law decay. We found that unit timescales within an LSTM, which are determined by the forget gate bias, should follow an Inverse Gamma distribution. Experiments then showed that LSTM language models trained on natural English text learn to approximate this theoretical distribution. Further, we found that explicitly imposing the theoretical distribution upon the model during training yielded better language model perplexity overall, with particular improvements for predicting low-frequency (rare) words. Moreover, the explicit multi-timescale model selectively routes information about different types of words through units with different timescales, potentially improving model interpretability. These results demonstrate the importance of careful, theoretically-motivated analysis of memory and timescale in language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69640430-af3a-45d4-a3c9-78992e25ecbcCited by top-tier papers9
- Interpretable multi-timescale models for predicting fMRI responses to continuous natural speechShailee Jain, Vy A. Vo, Shivangi Mahto, Amanda LeBel et al.NeurIPS 2020 · 58 citations
- Understanding Adaptive, Multiscale Temporal Integration In Deep Speech Recognition SystemsMenoua Keshishian, Samuel Norman-Haignere, Nima MesgaraniNeurIPS 2021 · 16 citations
- Emergent mechanisms for long timescales depend on training curriculum and affect performance in memory tasksSina Khajehabdollahi, Roxana Zeraati, Emmanouil Giannakakis, Tim Jakob Schäfer et al.ICLR 2024 · 8 citations
- Large language models transition from integrating across position-yoked, exponential windows to structure-yoked, power-law windowsDavid Skrill, Samuel Norman-HaignereNeurIPS 2023 · 7 citations
- The Expressive Leaky Memory Neuron: an Efficient and Expressive Phenomenological Neuron Model Can Solve Long-Horizon TasksAaron Spieler, Nasim Rahaman, Georg Martius, Bernhard Schölkopf et al.ICLR 2024 · 7 citations
Builds on4
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
- Interpretable multi-timescale models for predicting fMRI responses to continuous natural speechShailee Jain, Vy A. Vo, Shivangi Mahto, Amanda LeBel et al.NeurIPS 2020 · 58 citations
- Approximating Stacked and Bidirectional Recurrent Architectures with the Delayed Recurrent Neural NetworkJavier Turek, Shailee Jain, Vy A. Vo, Mihai Capota et al.ICML 2020 · 13 citations
- Mogrifier LSTMGábor Melis, Tomás Kociský, Phil BlunsomICLR 2020
Related papers
- Bivariate Beta-LSTMKyungwoo Song, JoonHo Jang, Seungjae Shin, Il-Chul MoonAAAI 2020 · 6 citations
- xLSTM: Extended Long Short-Term MemoryMaximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer et al.NeurIPS 2024 · 703 citations
- Deriving Neural Scaling Laws from the Statistics of Natural LanguageFrancesco Cagnetta, Allan Raventos, Surya Ganguli, Matthieu WyartICML 2026
- Unlocking the Power of LSTM for Long Term Time Series ForecastingYaxuan Kong, Zepu Wang, Yuqi Nie, Tian Zhou et al.AAAI 2025 · 92 citations
- Do RNN and LSTM have Long Memory?Jingyu Zhao, Feiqing Huang, Jia Lv, Yanjie Duan et al.ICML 2020 · 184 citations
