projUNN: efficient method for training deep networks with unitary matrices
Bobak Toussi Kiani, Randall Balestriero, Yann LeCun, Seth Lloyd
Abstract
In learning with recurrent or very deep feed-forward networks, employing unitary matrices in each layer can be very effective at maintaining long-range stability. However, restricting network parameters to be unitary typically comes at the cost of expensive parameterizations or increased training runtime. We propose instead an efficient method based on rank- updates -- or their rank- approximation -- that maintains performance at a nearly optimal training runtime. We introduce two variants of this method, named Direct (projUNN-D) and Tangent (projUNN-T) projected Unitary Neural Networks, that can parameterize full -dimensional unitary or orthogonal matrices with a training runtime scaling as . Our method either projects low-rank gradients onto the closest unitary matrix (projUNN-T) or transports unitary matrices in the direction of the low-rank gradient (projUNN-D). Even in the fastest setting (), projUNN is able to train a model's unitary parameters to reach comparable performances against baseline implementations. In recurrent neural network settings, projUNN closely matches or exceeds benchmarked results from prior unitary neural networks. Finally, we preliminarily explore projUNN in training orthogonal convolutional neural networks, which are currently unable to outperform state of the art models but can potentially enhance stability and robustness at large depth.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59dd9fef-7287-4e19-a4bf-d037a8295320Cited by top-tier papers10
- Self-Supervised Learning with Lie Symmetries for Partial Differential EquationsGrégoire Mialon, Quentin Garrido, Hannah Lawrence, Danyal Rehman et al.NeurIPS 2023 · 33 citations
- Stabilized Neural Differential Equations for Learning Dynamics with Explicit ConstraintsAlistair White, Niki Kilbertus, Maximilian Gelbrecht, Niklas BoersNeurIPS 2023 · 22 citations
- Unitary Convolutions for Learning on Graphs and GroupsBobak T. Kiani, Lukas Fesser, Melanie WeberNeurIPS 2024 · 13 citations
- Improved techniques for deterministic l2 robustnessSahil Singla, Soheil FeiziNeurIPS 2022 · 13 citations
- Generating Universal Adversarial Perturbations for Quantum ClassifiersGautham Anil, Vishnu Vinod, Apurva NarayanAAAI 2024 · 10 citations
Builds on11
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Quantum Algorithms for Deep Convolutional Neural NetworksIordanis Kerenidis, Jonas Landman, Anupam PrakashICLR 2020 · 163 citations
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley TransformJun Li, Fuxin Li, Sinisa TodorovicICLR 2020 · 139 citations
- Orthogonalizing Convolutional Layers with the Cayley TransformAsher Trockman, J. Zico KolterICLR 2021 · 137 citations
- Skew Orthogonal ConvolutionsSahil Singla, Soheil FeiziICML 2021 · 76 citations
Related papers
- Coordinate Descent on the Orthogonal Group for Recurrent Neural Network TrainingEstelle M. Massart, Vinayak AbrolAAAI 2022 · 13 citations
- Approximated Orthogonal Projection Unit: Stabilizing Regression Network Training Using Natural GradientShaoqi Wang, Chunjie Yang, Siwei LouNeurIPS 2024 · 1 citation
- Finite-Width Neural Tangent Kernels from Feynman DiagramsMax Guillen, Philipp Misof, Jan GerkenICML 2026 · 1 citation
- Scaling-up Diverse Orthogonal Convolutional Networks by a Paraunitary FrameworkJiahao Su, Wonmin Byeon, Furong HuangICML 2022 · 1 citation
- Controllable Orthogonalization in Training DNNsLei Huang, Li Liu, Fan Zhu, Diwen Wan et al.CVPR 2020
