Fast Differentiable Matrix Square Root
Yue Song, Nicu Sebe, Wei Wang
摘要
Computing the matrix square root or its inverse in a differentiable manner is important in a variety of computer vision tasks. Previous methods either adopt the Singular Value Decomposition (SVD) to explicitly factorize the matrix or use the Newton-Schulz iteration (NS iteration) to derive the approximate solution. However, both methods are not computationally efficient enough in either the forward pass or in the backward pass. In this paper, we propose two more efficient variants to compute the differentiable matrix square root. For the forward propagation, one method is to use Matrix Taylor Polynomial (MTP), and the other method is to use Matrix Padé Approximants (MPA). The backward gradient is computed by iteratively solving the continuous-time Lyapunov equation using the matrix sign function. Both methods yield considerable speed-up compared with the SVD or the Newton-Schulz iteration. Experimental results on the de-correlated batch normalization and second-order vision transformer demonstrate that our methods can also achieve competitive and even slightly better performances. The code is available at https://github.com/KingJamesSong/FastDifferentiableMatSqrt.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- DropCov: A Simple yet Effective Method for Improving Deep ArchitecturesQilong Wang, Mingze Gao, Zhaolin Zhang, Jiangtao Xie 等NeurIPS 2022 · 被引用 13 次
- Grouping Matrix Based Graph Pooling with Adaptive Number of ClustersSung Moon Ko, Sungjun Cho, Dae-Woong Jeong, Sehui Han 等AAAI 2023 · 被引用 12 次
- Enhancing LLM Training via Spectral ClippingXiaowen Jiang, Andrei Semenov, Sebastian StichICML 2026 · 被引用 4 次
- MuonSSM: Orthogonalizing State Space Models for Sequence ModelingThai Khanh Nguyen, Uyen N.B. Vo, Thieu Vo, Tan Nguyen 等ICML 2026 · 被引用 1 次
- Understanding Matrix Function Normalizations in Covariance Pooling through the Lens of Riemannian GeometryZiheng Chen, Yue Song, Xiaojun Wu, Gaowen Liu 等ICLR 2025 · 被引用 1 次
它引用的顶会 Paper8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Whitening for Self-Supervised Representation LearningAleksandr Ermolov, Aliaksandr Siarohin, Enver Sangineto, Nicu SebeICML 2021 · 被引用 378 次
- Switchable Whitening for Deep Representation LearningXingang Pan, Xiaohang Zhan, Jianping Shi, Xiaoou Tang 等ICCV 2019 · 被引用 204 次
- Why Approximate Matrix Square Root Outperforms Accurate SVD in Global Covariance Pooling?Yue Song, Nicu Sebe, Wei WangICCV 2021 · 被引用 39 次
- An Investigation Into the Stochasticity of Batch WhiteningLei Huang, Lei Zhao, Yi Zhou, Fan Zhu 等CVPR 2020
相关 Paper
- SPAN: A Stochastic Projected Approximate Newton MethodXunpeng Huang, Xianfeng Liang, Zhengyang Liu, Lei Li 等AAAI 2020 · 被引用 4 次
- QT-ViT: Improving Linear Attention in ViT with Quadratic Taylor ExpansionYixing Xu, Chao Li, Dong Li, Xiao Sheng 等NeurIPS 2024 · 被引用 7 次
- What if Neural Networks had SVDs?Alexander Mathiasen, Frederik Hvilshøj, Jakob Rødsgaard Jørgensen, Anshul Nasery 等NeurIPS 2020 · 被引用 12 次
- Collapsing Taylor Mode Automatic DifferentiationFelix Dangel, Tim Siebert, Marius Zeinhofer, Andrea WaltherNeurIPS 2025 · 被引用 1 次
- SOFT: Softmax-free Transformer with Linear ComplexityJiachen Lu, Jinghan Yao, Junge Zhang, Xiatian Zhu 等NeurIPS 2021 · 被引用 232 次
