Understanding the Covariance Structure of Convolutional Filters
Asher Trockman, Devin Willmott, J. Zico Kolter
Abstract
Neural network weights are typically initialized at random from univariate distributions, controlling just the variance of individual weights even in highlystructured operations like convolutions. Recent ViT-inspired convolutional networks such as ConvMixer and ConvNeXt use large-kernel depthwise convolutions whose learned filters have notable structure; this presents an opportunity to study their empirical covariances. In this work, we first observe that such learned filters have highly-structured covariance matrices, and moreover, we find that covariances calculated from small networks may be used to effectively initialize a variety of larger networks of different depths, widths, patch sizes, and kernel sizes, indicating a degree of model-independence to the covariance structure. Motivated by these findings, we then propose a learning-free multivariate initialization scheme for convolutional filters using a simple, closed-form construction of their covariance. Models using our initialization outperform those using traditional univariate initializations, and typically meet or exceed the performance of those initialized from the covariances of learned filters; in some cases, this improvement can be achieved without training the depthwise convolutional filters at all.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e2f07da-9cc5-4a94-856c-3c82b068c974Cited by top-tier papers7
- White-Box Transformers via Sparse Rate ReductionYaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu et al.NeurIPS 2023 · 149 citations
- Mimetic Initialization of Self-Attention LayersAsher Trockman, J. Zico KolterICML 2023 · 54 citations
- Initializing Models with Larger OnesZhiqiu Xu, Yanjie Chen, Kirill Vishniakov, Yida Yin et al.ICLR 2024 · 40 citations
- Structured Initialization for Vision TransformersJianqiao Zheng, Xueqian Li, Hemanth Saratchandran, Simon LuceyNeurIPS 2025 · 6 citations
- The Quest for Universal Master Key Filters in DS-CNNsZahra Babaiee, Peyman M. Kiasari, Daniela Rus, Radu GrosuNeurIPS 2025 · 2 citations
Builds on4
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- On the Connection between Local Attention and Dynamic Depth-wise ConvolutionQi Han, Zejia Fan, Qi Dai, Lei Sun et al.ICLR 2022 · 144 citations
- More ConvNets in the 2020s: Scaling up Kernels Beyond 51x51 using SparsityShiwei Liu, Tianlong Chen, Xiaohan Chen, Xuxi Chen et al.ICLR 2023 · 87 citations
Related papers
- InceptionNeXt: When Inception Meets ConvNeXtWeihao Yu, Pan Zhou, Shuicheng Yan, Xinchao WangCVPR 2024 · 326 citations
- Beyond Signal Propagation: Is Feature Diversity Necessary in Deep Neural Network Initialization?Yaniv Blumenfeld, Dar Gilboa, Daniel SoudryICML 2020 · 18 citations
- Unveiling the Unseen: Identifiable Clusters in Trained Depthwise Convolutional KernelsZahra Babaiee, Peyman M. Kiasari, Daniela Rus, Radu GrosuICLR 2024 · 9 citations
- AutoInit: Analytic Signal-Preserving Weight Initialization for Neural NetworksGarrett Bingham, Risto MiikkulainenAAAI 2023 · 6 citations
- Principled Architecture-aware Scaling of HyperparametersWuyang Chen, Junru Wu, Zhangyang Wang, Boris HaninICLR 2024 · 3 citations
