Multi-timescale Representation Learning in LSTM Language Models
Shivangi Mahto, Vy Ai Vo, Javier S. Turek, Alexander Huth
摘要
Language models must capture statistical dependencies between words at timescales ranging from very short to very long. Earlier work has demonstrated that dependencies in natural language tend to decay with distance between words according to a power law. However, it is unclear how this knowledge can be used for analyzing or designing neural network language models. In this work, we derived a theory for how the memory gating mechanism in long short-term memory (LSTM) language models can capture power law decay. We found that unit timescales within an LSTM, which are determined by the forget gate bias, should follow an Inverse Gamma distribution. Experiments then showed that LSTM language models trained on natural English text learn to approximate this theoretical distribution. Further, we found that explicitly imposing the theoretical distribution upon the model during training yielded better language model perplexity overall, with particular improvements for predicting low-frequency (rare) words. Moreover, the explicit multi-timescale model selectively routes information about different types of words through units with different timescales, potentially improving model interpretability. These results demonstrate the importance of careful, theoretically-motivated analysis of memory and timescale in language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Interpretable multi-timescale models for predicting fMRI responses to continuous natural speechShailee Jain, Vy A. Vo, Shivangi Mahto, Amanda LeBel 等NeurIPS 2020 · 被引用 58 次
- Understanding Adaptive, Multiscale Temporal Integration In Deep Speech Recognition SystemsMenoua Keshishian, Samuel Norman-Haignere, Nima MesgaraniNeurIPS 2021 · 被引用 16 次
- Emergent mechanisms for long timescales depend on training curriculum and affect performance in memory tasksSina Khajehabdollahi, Roxana Zeraati, Emmanouil Giannakakis, Tim Jakob Schäfer 等ICLR 2024 · 被引用 8 次
- Large language models transition from integrating across position-yoked, exponential windows to structure-yoked, power-law windowsDavid Skrill, Samuel Norman-HaignereNeurIPS 2023 · 被引用 7 次
- The Expressive Leaky Memory Neuron: an Efficient and Expressive Phenomenological Neuron Model Can Solve Long-Horizon TasksAaron Spieler, Nasim Rahaman, Georg Martius, Bernhard Schölkopf 等ICLR 2024 · 被引用 7 次
它引用的顶会 Paper4
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 被引用 629 次
- Interpretable multi-timescale models for predicting fMRI responses to continuous natural speechShailee Jain, Vy A. Vo, Shivangi Mahto, Amanda LeBel 等NeurIPS 2020 · 被引用 58 次
- Approximating Stacked and Bidirectional Recurrent Architectures with the Delayed Recurrent Neural NetworkJavier Turek, Shailee Jain, Vy A. Vo, Mihai Capota 等ICML 2020 · 被引用 13 次
- Mogrifier LSTMGábor Melis, Tomás Kociský, Phil BlunsomICLR 2020
相关 Paper
- Bivariate Beta-LSTMKyungwoo Song, JoonHo Jang, Seungjae Shin, Il-Chul MoonAAAI 2020 · 被引用 6 次
- xLSTM: Extended Long Short-Term MemoryMaximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer 等NeurIPS 2024 · 被引用 703 次
- Deriving Neural Scaling Laws from the Statistics of Natural LanguageFrancesco Cagnetta, Allan Raventos, Surya Ganguli, Matthieu WyartICML 2026
- Unlocking the Power of LSTM for Long Term Time Series ForecastingYaxuan Kong, Zepu Wang, Yuqi Nie, Tian Zhou 等AAAI 2025 · 被引用 92 次
- Do RNN and LSTM have Long Memory?Jingyu Zhao, Feiqing Huang, Jia Lv, Yanjie Duan 等ICML 2020 · 被引用 184 次
