A Tensor Decomposition Perspective on Second-order RNNs
Maude Lizaire, Michael Rizvi-Martel, Marawan Gamal Abdel Hameed, Guillaume Rabusseau
摘要
Second-order Recurrent Neural Networks (2RNNs) extend RNNs by leveraging second-order interactions for sequence modelling. These models are provably more expressive than their first-order counterparts and have connections to well-studied models from formal language theory. However, their large parameter tensor makes computations intractable. To circumvent this issue, one approach known as MIRNN consists in limiting the type of interactions used by the model. Another is to leverage tensor decomposition to diminish the parameter count. In this work, we study the model resulting from parameterizing 2RNNs using the CP decomposition, which we call CPRNN. Intuitively, the rank of the decomposition should reduce expressivity. We analyze how rank and hidden size affect model capacity and show the relationships between RNNs, 2RNNs, MIRNNs, and CPRNNs based on these parameters. We support these results empirically with experiments on the Penn Treebank dataset which demonstrate that, with a fixed parameter budget, CPRNNs outperforms RNNs, 2RNNs, and MIRNNs with the right choice of rank and hidden size.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando 等ICML 2023 · 被引用 474 次
- Multiplicative Interactions and Where to Find ThemSiddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz 等ICLR 2020 · 被引用 152 次
- Neural Networks and the Chomsky HierarchyGrégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein 等ICLR 2023 · 被引用 45 次
- Scalable Interpretability via PolynomialsAbhimanyu Dubey, Filip Radenovic, Dhruv MahajanNeurIPS 2022 · 被引用 42 次
相关 Paper
- Compact Autoregressive NetworkDi Wang, Feiqing Huang, Jingyu Zhao, Guodong Li 等AAAI 2020 · 被引用 5 次
- High-Order Pooling for Graph Neural Networks with Tensor DecompositionChenqing Hua, Guillaume Rabusseau, Jian TangNeurIPS 2022 · 被引用 45 次
- Towards Extremely Compact RNNs for Video Recognition With Fully Decomposed Hierarchical Tucker StructureMiao Yin, Siyu Liao, Xiao-Yang Liu, Xiaodong Wang 等CVPR 2021
- Decomposing Temporal High-Order Interactions via Latent ODEsShibo Li, Robert M. Kirby, Shandian ZheICML 2022 · 被引用 6 次
- Towards Efficient Tensor Decomposition-Based DNN Model Compression With Optimization FrameworkMiao Yin, Yang Sui, Siyu Liao, Bo YuanCVPR 2021
