Revitalizing SVD for Global Covariance Pooling: Halley's Method to Overcome Over-Flattening
Jiawei Gu, Ziyue Qiao, Xinming Li, Zechao Li
摘要
Global Covariance Pooling (GCP) has garnered increasing attention in visual recognition tasks, where second-order statistics frequently yield stronger representations than first-order approaches. However, two main streams of GCP-Newton-Schulzbased iSQRT-COV and exact or near-exact SVD methods-struggle at opposite ends of the training spectrum. While iSQRT-COV stabilizes early learning by avoiding large gradient explosions, it over-compresses significant eigenvalues in later stages, causing an over-flattening phenomenon that stalls final accuracy. In contrast, SVD-based methods excel at preserving the high-eigenvalue structure essential for deep networks but suffer from sensitivity to small eigenvalue gaps early on. We propose Halley-SVD, a high-order iterative method that unites the smooth gradient advantages of iSQRT-COV with the late-stage fidelity of SVD. Grounded in Halley's iteration, our approach obviates explicit divisions by (λ i -λ j ) and forgoes threshold-or polynomial-based heuristics. As a result, it prevents both early gradient explosions and the excessive compression of large eigenvalues. Extensive experiments on CNNs and transformer architectures show that Halley-SVD consistently and robustly outperforms iSQRT-COV at large model scales and batch sizes, achieving higher overall accuracy without mid-training switches or custom truncations. This work provides a new solution to the long-standing dichotomy in GCP, illustrating how high-order methods can balance robustness and spectral precision to fully harness the representational power of modern deep networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- IMAGPose: A Unified Conditional Framework for Pose-Guided Person GenerationFei Shen, Jinhui TangNeurIPS 2024 · 被引用 172 次
相关 Paper
- Why Approximate Matrix Square Root Outperforms Accurate SVD in Global Covariance Pooling?Yue Song, Nicu Sebe, Wei WangICCV 2021 · 被引用 39 次
- What Deep CNNs Benefit From Global Covariance Pooling: An Optimization PerspectiveQilong Wang, Li Zhang, Banggu Wu, Dongwei Ren 等CVPR 2020
- DropCov: A Simple yet Effective Method for Improving Deep ArchitecturesQilong Wang, Mingze Gao, Zhaolin Zhang, Jiangtao Xie 等NeurIPS 2022 · 被引用 13 次
- Understanding Matrix Function Normalizations in Covariance Pooling through the Lens of Riemannian GeometryZiheng Chen, Yue Song, Xiaojun Wu, Gaowen Liu 等ICLR 2025 · 被引用 1 次
- DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root SolversIonut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan 等ICML 2026 · 被引用 1 次
