Understanding the Covariance Structure of Convolutional Filters
Asher Trockman, Devin Willmott, J. Zico Kolter
摘要
Neural network weights are typically initialized at random from univariate distributions, controlling just the variance of individual weights even in highlystructured operations like convolutions. Recent ViT-inspired convolutional networks such as ConvMixer and ConvNeXt use large-kernel depthwise convolutions whose learned filters have notable structure; this presents an opportunity to study their empirical covariances. In this work, we first observe that such learned filters have highly-structured covariance matrices, and moreover, we find that covariances calculated from small networks may be used to effectively initialize a variety of larger networks of different depths, widths, patch sizes, and kernel sizes, indicating a degree of model-independence to the covariance structure. Motivated by these findings, we then propose a learning-free multivariate initialization scheme for convolutional filters using a simple, closed-form construction of their covariance. Models using our initialization outperform those using traditional univariate initializations, and typically meet or exceed the performance of those initialized from the covariances of learned filters; in some cases, this improvement can be achieved without training the depthwise convolutional filters at all.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- White-Box Transformers via Sparse Rate ReductionYaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu 等NeurIPS 2023 · 被引用 149 次
- Mimetic Initialization of Self-Attention LayersAsher Trockman, J. Zico KolterICML 2023 · 被引用 54 次
- Initializing Models with Larger OnesZhiqiu Xu, Yanjie Chen, Kirill Vishniakov, Yida Yin 等ICLR 2024 · 被引用 40 次
- Structured Initialization for Vision TransformersJianqiao Zheng, Xueqian Li, Hemanth Saratchandran, Simon LuceyNeurIPS 2025 · 被引用 6 次
- The Quest for Universal Master Key Filters in DS-CNNsZahra Babaiee, Peyman M. Kiasari, Daniela Rus, Radu GrosuNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper4
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- On the Connection between Local Attention and Dynamic Depth-wise ConvolutionQi Han, Zejia Fan, Qi Dai, Lei Sun 等ICLR 2022 · 被引用 144 次
- More ConvNets in the 2020s: Scaling up Kernels Beyond 51x51 using SparsityShiwei Liu, Tianlong Chen, Xiaohan Chen, Xuxi Chen 等ICLR 2023 · 被引用 87 次
相关 Paper
- InceptionNeXt: When Inception Meets ConvNeXtWeihao Yu, Pan Zhou, Shuicheng Yan, Xinchao WangCVPR 2024 · 被引用 326 次
- Beyond Signal Propagation: Is Feature Diversity Necessary in Deep Neural Network Initialization?Yaniv Blumenfeld, Dar Gilboa, Daniel SoudryICML 2020 · 被引用 18 次
- Unveiling the Unseen: Identifiable Clusters in Trained Depthwise Convolutional KernelsZahra Babaiee, Peyman M. Kiasari, Daniela Rus, Radu GrosuICLR 2024 · 被引用 9 次
- AutoInit: Analytic Signal-Preserving Weight Initialization for Neural NetworksGarrett Bingham, Risto MiikkulainenAAAI 2023 · 被引用 6 次
- Principled Architecture-aware Scaling of HyperparametersWuyang Chen, Junru Wu, Zhangyang Wang, Boris HaninICLR 2024 · 被引用 3 次
