Fast Differentiable Matrix Square Root
Yue Song, Nicu Sebe, Wei Wang
Abstract
Computing the matrix square root or its inverse in a differentiable manner is important in a variety of computer vision tasks. Previous methods either adopt the Singular Value Decomposition (SVD) to explicitly factorize the matrix or use the Newton-Schulz iteration (NS iteration) to derive the approximate solution. However, both methods are not computationally efficient enough in either the forward pass or in the backward pass. In this paper, we propose two more efficient variants to compute the differentiable matrix square root. For the forward propagation, one method is to use Matrix Taylor Polynomial (MTP), and the other method is to use Matrix Padé Approximants (MPA). The backward gradient is computed by iteratively solving the continuous-time Lyapunov equation using the matrix sign function. Both methods yield considerable speed-up compared with the SVD or the Newton-Schulz iteration. Experimental results on the de-correlated batch normalization and second-order vision transformer demonstrate that our methods can also achieve competitive and even slightly better performances. The code is available at https://github.com/KingJamesSong/FastDifferentiableMatSqrt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- DropCov: A Simple yet Effective Method for Improving Deep ArchitecturesQilong Wang, Mingze Gao, Zhaolin Zhang, Jiangtao Xie et al.NeurIPS 2022 · 13 citations
- Grouping Matrix Based Graph Pooling with Adaptive Number of ClustersSung Moon Ko, Sungjun Cho, Dae-Woong Jeong, Sehui Han et al.AAAI 2023 · 12 citations
- Enhancing LLM Training via Spectral ClippingXiaowen Jiang, Andrei Semenov, Sebastian StichICML 2026 · 4 citations
- MuonSSM: Orthogonalizing State Space Models for Sequence ModelingThai Khanh Nguyen, Uyen N.B. Vo, Thieu Vo, Tan Nguyen et al.ICML 2026 · 1 citation
- Understanding Matrix Function Normalizations in Covariance Pooling through the Lens of Riemannian GeometryZiheng Chen, Yue Song, Xiaojun Wu, Gaowen Liu et al.ICLR 2025 · 1 citation
Builds on8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Whitening for Self-Supervised Representation LearningAleksandr Ermolov, Aliaksandr Siarohin, Enver Sangineto, Nicu SebeICML 2021 · 378 citations
- Switchable Whitening for Deep Representation LearningXingang Pan, Xiaohang Zhan, Jianping Shi, Xiaoou Tang et al.ICCV 2019 · 204 citations
- Why Approximate Matrix Square Root Outperforms Accurate SVD in Global Covariance Pooling?Yue Song, Nicu Sebe, Wei WangICCV 2021 · 39 citations
- An Investigation Into the Stochasticity of Batch WhiteningLei Huang, Lei Zhao, Yi Zhou, Fan Zhu et al.CVPR 2020
Related papers
- SPAN: A Stochastic Projected Approximate Newton MethodXunpeng Huang, Xianfeng Liang, Zhengyang Liu, Lei Li et al.AAAI 2020 · 4 citations
- QT-ViT: Improving Linear Attention in ViT with Quadratic Taylor ExpansionYixing Xu, Chao Li, Dong Li, Xiao Sheng et al.NeurIPS 2024 · 7 citations
- What if Neural Networks had SVDs?Alexander Mathiasen, Frederik Hvilshøj, Jakob Rødsgaard Jørgensen, Anshul Nasery et al.NeurIPS 2020 · 12 citations
- Collapsing Taylor Mode Automatic DifferentiationFelix Dangel, Tim Siebert, Marius Zeinhofer, Andrea WaltherNeurIPS 2025 · 1 citation
- SOFT: Softmax-free Transformer with Linear ComplexityJiachen Lu, Jinghan Yao, Junge Zhang, Xiatian Zhu et al.NeurIPS 2021 · 232 citations
