Why Approximate Matrix Square Root Outperforms Accurate SVD in Global Covariance Pooling?
Yue Song, Nicu Sebe, Wei Wang
Abstract
Global Covariance Pooling (GCP) aims at exploiting the second-order statistics of the convolutional feature. Its effectiveness has been demonstrated in boosting the classification performance of Convolutional Neural Networks (CNNs). Singular Value Decomposition (SVD) is used in GCP to compute the matrix square root. However, the approximate matrix square root calculated using Newton-Schulz iteration [14] outperforms the accurate one computed via SVD [15]. We empirically analyze the reason behind the performance gap from the perspectives of data precision and gradient smoothness. Various remedies for computing smooth SVD gradients are investigated. Based on our observation and analyses, a hybrid training protocol is proposed for SVD-based GCP meta-layers such that competitive performances can be achieved against Newton-Schulz iteration. Moreover, we propose a new GCP meta-layer that uses SVD in the forward pass, and Padé approximants in the backward propagation to compute the gradients. The proposed meta-layer has been integrated into different CNN models and achieves state-of-the-art performances on both large-scale and fine-grained datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1cf468d5-3bb8-40e7-a670-1cc513d34e9eCited by top-tier papers12
- Fast Differentiable Matrix Square RootYue Song, Nicu Sebe, Wei WangICLR 2022 · 18 citations
- RMLR: Extending Multinomial Logistic Regression into General GeometriesZiheng Chen, Yue Song, Rui Wang, Xiaojun Wu et al.NeurIPS 2024 · 16 citations
- DropCov: A Simple yet Effective Method for Improving Deep ArchitecturesQilong Wang, Mingze Gao, Zhaolin Zhang, Jiangtao Xie et al.NeurIPS 2022 · 13 citations
- Self-Consistency Training for Density-Functional-Theory Hamiltonian PredictionHe Zhang, Chang Liu, Zun Wang, Xinran Wei et al.ICML 2024 · 13 citations
- Grouping Matrix Based Graph Pooling with Adaptive Number of ClustersSung Moon Ko, Sungjun Cho, Dae-Woong Jeong, Sehui Han et al.AAAI 2023 · 12 citations
Builds on3
- Understanding Generalized Whitening and Coloring Transform for Universal Style TransferTai-Yin ChiuICCV 2019 · 32 citations
- An Investigation Into the Stochasticity of Batch WhiteningLei Huang, Lei Zhao, Yi Zhou, Fan Zhu et al.CVPR 2020
- What Deep CNNs Benefit From Global Covariance Pooling: An Optimization PerspectiveQilong Wang, Li Zhang, Banggu Wu, Dongwei Ren et al.CVPR 2020
Related papers
- Revitalizing SVD for Global Covariance Pooling: Halley's Method to Overcome Over-FlatteningJiawei Gu, Ziyue Qiao, Xinming Li, Zechao LiNeurIPS 2025 · 5 citations
- Understanding Matrix Function Normalizations in Covariance Pooling through the Lens of Riemannian GeometryZiheng Chen, Yue Song, Xiaojun Wu, Gaowen Liu et al.ICLR 2025 · 1 citation
- DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root SolversIonut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan et al.ICML 2026 · 1 citation
- A New Perspective on Shampoo's PreconditionerDepen Morwani, Itai Shapira, Nikhil Vyas, Eran Malach et al.ICLR 2025
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li et al.CVPR 2021
