projUNN: efficient method for training deep networks with unitary matrices
Bobak Toussi Kiani, Randall Balestriero, Yann LeCun, Seth Lloyd
摘要
In learning with recurrent or very deep feed-forward networks, employing unitary matrices in each layer can be very effective at maintaining long-range stability. However, restricting network parameters to be unitary typically comes at the cost of expensive parameterizations or increased training runtime. We propose instead an efficient method based on rank- updates -- or their rank- approximation -- that maintains performance at a nearly optimal training runtime. We introduce two variants of this method, named Direct (projUNN-D) and Tangent (projUNN-T) projected Unitary Neural Networks, that can parameterize full -dimensional unitary or orthogonal matrices with a training runtime scaling as . Our method either projects low-rank gradients onto the closest unitary matrix (projUNN-T) or transports unitary matrices in the direction of the low-rank gradient (projUNN-D). Even in the fastest setting (), projUNN is able to train a model's unitary parameters to reach comparable performances against baseline implementations. In recurrent neural network settings, projUNN closely matches or exceeds benchmarked results from prior unitary neural networks. Finally, we preliminarily explore projUNN in training orthogonal convolutional neural networks, which are currently unable to outperform state of the art models but can potentially enhance stability and robustness at large depth.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Self-Supervised Learning with Lie Symmetries for Partial Differential EquationsGrégoire Mialon, Quentin Garrido, Hannah Lawrence, Danyal Rehman 等NeurIPS 2023 · 被引用 33 次
- Stabilized Neural Differential Equations for Learning Dynamics with Explicit ConstraintsAlistair White, Niki Kilbertus, Maximilian Gelbrecht, Niklas BoersNeurIPS 2023 · 被引用 22 次
- Unitary Convolutions for Learning on Graphs and GroupsBobak T. Kiani, Lukas Fesser, Melanie WeberNeurIPS 2024 · 被引用 13 次
- Improved techniques for deterministic l2 robustnessSahil Singla, Soheil FeiziNeurIPS 2022 · 被引用 13 次
- Generating Universal Adversarial Perturbations for Quantum ClassifiersGautham Anil, Vishnu Vinod, Apurva NarayanAAAI 2024 · 被引用 10 次
它引用的顶会 Paper11
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Quantum Algorithms for Deep Convolutional Neural NetworksIordanis Kerenidis, Jonas Landman, Anupam PrakashICLR 2020 · 被引用 163 次
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley TransformJun Li, Fuxin Li, Sinisa TodorovicICLR 2020 · 被引用 139 次
- Orthogonalizing Convolutional Layers with the Cayley TransformAsher Trockman, J. Zico KolterICLR 2021 · 被引用 137 次
- Skew Orthogonal ConvolutionsSahil Singla, Soheil FeiziICML 2021 · 被引用 76 次
相关 Paper
- Coordinate Descent on the Orthogonal Group for Recurrent Neural Network TrainingEstelle M. Massart, Vinayak AbrolAAAI 2022 · 被引用 13 次
- Approximated Orthogonal Projection Unit: Stabilizing Regression Network Training Using Natural GradientShaoqi Wang, Chunjie Yang, Siwei LouNeurIPS 2024 · 被引用 1 次
- Finite-Width Neural Tangent Kernels from Feynman DiagramsMax Guillen, Philipp Misof, Jan GerkenICML 2026 · 被引用 1 次
- Scaling-up Diverse Orthogonal Convolutional Networks by a Paraunitary FrameworkJiahao Su, Wonmin Byeon, Furong HuangICML 2022 · 被引用 1 次
- Controllable Orthogonalization in Training DNNsLei Huang, Li Liu, Fan Zhu, Diwen Wan 等CVPR 2020
