Neural Dependencies Emerging from Learning Massive Categories
Ruili Feng, Kecheng Zheng, Kai Zhu, Yujun Shen, Jian Zhao, Yukun Huang, Deli Zhao, Jingren Zhou, Michael I. Jordan, Zheng-Jun Zha
Abstract
This work presents two astonishing findings on neural networks learned for large-scale image classification. 1) Given a well-trained model, the logits predicted for some category can be directly obtained by linearly combining the predictions of a few other categories, which we call neural dependency. 2) Neural dependencies exist not only within a single model, but even between two independently learned models, regardless of their architectures. Towards a theoretical analysis of such phenomena, we demonstrate that identifying neural dependencies is equivalent to solving the Covariance Lasso (CovLasso) regression problem proposed in this paper. Through investigating the properties of the problem solution, we confirm that neural dependency is guaranteed by a redundant logit covariance matrix, which condition is easily met given massive categories, and that neural dependency is highly sparse, implying that one category correlates to only a few others. We further empirically show the potential of neural dependencies in understanding internal data correlations, generalizing models to unseen categories, and improving model robustness with a dependency-derived regularizer. Code to reproduce the results in this paper is available at https://github.com/RuiLiFengiNeural-Dependencies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf672eff-3db6-4121-9600-118ddffe437cBuilds on6
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Rank Diminishing in Deep Neural NetworksRuili Feng, Kecheng Zheng, Yukun Huang, Deli Zhao et al.NeurIPS 2022 · 64 citations
- Towards Understanding Hierarchical Learning: Benefits of Neural RepresentationsMinshuo Chen, Yu Bai, Jason D. Lee, Tuo Zhao et al.NeurIPS 2020 · 61 citations
Related papers
- Correlated Input-Dependent Label Noise in Large-Scale Image ClassificationMark Collier, Basil Mustafa, Efi Kokiopoulou, Rodolphe Jenatton et al.CVPR 2021
- The Prevalence of Neural Collapse in Neural Multivariate RegressionGeorge Andriopoulos, Zixuan Dong, Li Guo, Zifan Zhao et al.NeurIPS 2024 · 24 citations
- Conditional Temporal Neural Processes with Covariance LossBoseon Yoo, Jiwoo Lee, Janghoon Ju, Seijun Chung et al.ICML 2021 · 19 citations
- Measuring Dependence with Matrix-based Entropy FunctionalShujian Yu, Francesco Alesiani, Xi Yu, Robert Jenssen et al.AAAI 2021 · 36 citations
- Efficient Conditionally Invariant Representation LearningRoman Pogodin, Namrata Deka, Yazhe Li, Danica J. Sutherland et al.ICLR 2023 · 2 citations
