Why Approximate Matrix Square Root Outperforms Accurate SVD in Global Covariance Pooling?
Yue Song, Nicu Sebe, Wei Wang
摘要
Global Covariance Pooling (GCP) aims at exploiting the second-order statistics of the convolutional feature. Its effectiveness has been demonstrated in boosting the classification performance of Convolutional Neural Networks (CNNs). Singular Value Decomposition (SVD) is used in GCP to compute the matrix square root. However, the approximate matrix square root calculated using Newton-Schulz iteration [14] outperforms the accurate one computed via SVD [15]. We empirically analyze the reason behind the performance gap from the perspectives of data precision and gradient smoothness. Various remedies for computing smooth SVD gradients are investigated. Based on our observation and analyses, a hybrid training protocol is proposed for SVD-based GCP meta-layers such that competitive performances can be achieved against Newton-Schulz iteration. Moreover, we propose a new GCP meta-layer that uses SVD in the forward pass, and Padé approximants in the backward propagation to compute the gradients. The proposed meta-layer has been integrated into different CNN models and achieves state-of-the-art performances on both large-scale and fine-grained datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Fast Differentiable Matrix Square RootYue Song, Nicu Sebe, Wei WangICLR 2022 · 被引用 18 次
- RMLR: Extending Multinomial Logistic Regression into General GeometriesZiheng Chen, Yue Song, Rui Wang, Xiaojun Wu 等NeurIPS 2024 · 被引用 16 次
- DropCov: A Simple yet Effective Method for Improving Deep ArchitecturesQilong Wang, Mingze Gao, Zhaolin Zhang, Jiangtao Xie 等NeurIPS 2022 · 被引用 13 次
- Self-Consistency Training for Density-Functional-Theory Hamiltonian PredictionHe Zhang, Chang Liu, Zun Wang, Xinran Wei 等ICML 2024 · 被引用 13 次
- Grouping Matrix Based Graph Pooling with Adaptive Number of ClustersSung Moon Ko, Sungjun Cho, Dae-Woong Jeong, Sehui Han 等AAAI 2023 · 被引用 12 次
它引用的顶会 Paper3
- Understanding Generalized Whitening and Coloring Transform for Universal Style TransferTai-Yin ChiuICCV 2019 · 被引用 32 次
- An Investigation Into the Stochasticity of Batch WhiteningLei Huang, Lei Zhao, Yi Zhou, Fan Zhu 等CVPR 2020
- What Deep CNNs Benefit From Global Covariance Pooling: An Optimization PerspectiveQilong Wang, Li Zhang, Banggu Wu, Dongwei Ren 等CVPR 2020
相关 Paper
- Revitalizing SVD for Global Covariance Pooling: Halley's Method to Overcome Over-FlatteningJiawei Gu, Ziyue Qiao, Xinming Li, Zechao LiNeurIPS 2025 · 被引用 5 次
- Understanding Matrix Function Normalizations in Covariance Pooling through the Lens of Riemannian GeometryZiheng Chen, Yue Song, Xiaojun Wu, Gaowen Liu 等ICLR 2025 · 被引用 1 次
- DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root SolversIonut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan 等ICML 2026 · 被引用 1 次
- A New Perspective on Shampoo's PreconditionerDepen Morwani, Itai Shapira, Nikhil Vyas, Eran Malach 等ICLR 2025
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li 等CVPR 2021
