Low-Rank Constraints for Fast Inference in Structured Models
Justin T. Chiu, Yuntian Deng, Alexander M. Rush
摘要
Structured distributions, i.e. distributions over combinatorial spaces, are commonly used to learn latent probabilistic representations from observed data. However, scaling these models is bottlenecked by the high computational and memory complexity with respect to the size of the latent representations. Common models such as Hidden Markov Models (HMMs) and Probabilistic Context-Free Grammars (PCFGs) require time and space quadratic and cubic in the number of hidden states respectively. This work demonstrates a simple approach to reduce the computational and memory complexity of a large class of structured models. We show that by viewing the central inference step as a matrix-vector product and using a low-rank constraint, we can trade off model expressivity and speed via the rank. Experiments with neural parameterized structured models for language modeling, polyphonic music modeling, unsupervised grammar induction, and video modeling show that our approach matches the accuracy of standard models at large state spaces while providing practical speedups. * Equal contribution Code is available here. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Rank Diminishing in Deep Neural NetworksRuili Feng, Kecheng Zheng, Yukun Huang, Deli Zhao 等NeurIPS 2022 · 被引用 64 次
- Learning to Grow Pretrained Models for Efficient Transformer TrainingPeihao Wang, Rameswar Panda, Lucas Torroba Hennigen, Philip Greengard 等ICLR 2023 · 被引用 13 次
- Scaling Structured Inference with RandomizationYao Fu, John P. Cunningham, Mirella LapataICML 2022 · 被引用 2 次
- Unsupervised Discontinuous Constituency Parsing with Mildly Context-Sensitive GrammarsSonglin Yang, Roger Levy, Yoon KimACL 2023 · 被引用 1 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Random Feature AttentionHao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz 等ICLR 2021 · 被引用 425 次
- Scaling Autoregressive Video ModelsDirk Weissenborn, Oscar Täckström, Jakob UszkoreitICLR 2020 · 被引用 252 次
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 等ICLR 2021 · 被引用 122 次
相关 Paper
- Probability Distribution Collapse: A Critical Bottleneck to Compact Unsupervised Neural Grammar InductionJinwook Park, Kangil KimEMNLP 2025
- eXponential FAmily Dynamical Systems (XFADS): Large-scale nonlinear Gaussian state-space modelingMatthew Dowling, Yuan Zhao, Il Memming ParkNeurIPS 2024 · 被引用 17 次
- Sequences of Logits Reveal the Low Rank Structure of Language ModelsNoah Golowich, Allen Liu, Abhishek ShettyICLR 2026 · 被引用 9 次
- Deep networks learn to parse uniform-depth context-free languages from local statisticsJack T. Parley, Francesco Cagnetta, Matthieu WyartICML 2026 · 被引用 4 次
- What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular LanguagesNadav Borenstein, Anej Svete, Robin Chan, Josef Valvoda 等ACL 2024
