Revitalizing SVD for Global Covariance Pooling: Halley's Method to Overcome Over-Flattening
Jiawei Gu, Ziyue Qiao, Xinming Li, Zechao Li
Abstract
Global Covariance Pooling (GCP) has garnered increasing attention in visual recognition tasks, where second-order statistics frequently yield stronger representations than first-order approaches. However, two main streams of GCP-Newton-Schulzbased iSQRT-COV and exact or near-exact SVD methods-struggle at opposite ends of the training spectrum. While iSQRT-COV stabilizes early learning by avoiding large gradient explosions, it over-compresses significant eigenvalues in later stages, causing an over-flattening phenomenon that stalls final accuracy. In contrast, SVD-based methods excel at preserving the high-eigenvalue structure essential for deep networks but suffer from sensitivity to small eigenvalue gaps early on. We propose Halley-SVD, a high-order iterative method that unites the smooth gradient advantages of iSQRT-COV with the late-stage fidelity of SVD. Grounded in Halley's iteration, our approach obviates explicit divisions by (λ i -λ j ) and forgoes threshold-or polynomial-based heuristics. As a result, it prevents both early gradient explosions and the excessive compression of large eigenvalues. Extensive experiments on CNNs and transformer architectures show that Halley-SVD consistently and robustly outperforms iSQRT-COV at large model scales and batch sizes, achieving higher overall accuracy without mid-training switches or custom truncations. This work provides a new solution to the long-standing dichotomy in GCP, illustrating how high-order methods can balance robustness and spectral precision to fully harness the representational power of modern deep networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- IMAGPose: A Unified Conditional Framework for Pose-Guided Person GenerationFei Shen, Jinhui TangNeurIPS 2024 · 172 citations
Related papers
- Why Approximate Matrix Square Root Outperforms Accurate SVD in Global Covariance Pooling?Yue Song, Nicu Sebe, Wei WangICCV 2021 · 39 citations
- What Deep CNNs Benefit From Global Covariance Pooling: An Optimization PerspectiveQilong Wang, Li Zhang, Banggu Wu, Dongwei Ren et al.CVPR 2020
- DropCov: A Simple yet Effective Method for Improving Deep ArchitecturesQilong Wang, Mingze Gao, Zhaolin Zhang, Jiangtao Xie et al.NeurIPS 2022 · 13 citations
- Understanding Matrix Function Normalizations in Covariance Pooling through the Lens of Riemannian GeometryZiheng Chen, Yue Song, Xiaojun Wu, Gaowen Liu et al.ICLR 2025 · 1 citation
- DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root SolversIonut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan et al.ICML 2026 · 1 citation
