Parameter Prediction for Unseen Deep Architectures
Boris Knyazev, Michal Drozdzal, Graham W. Taylor, Adriana Romero-Soriano
摘要
Deep learning has been successful in automating the design of features in machine learning pipelines. However, the algorithms optimizing neural network parameters remain largely hand-designed and computationally inefficient. We study if we can use deep learning to directly predict these parameters by exploiting the past knowledge of training other networks. We introduce a large-scale dataset of diverse computational graphs of neural architectures - DeepNets-1M - and use it to explore parameter prediction on CIFAR-10 and ImageNet. By leveraging advances in graph neural networks, we propose a hypernetwork that can predict performant parameters in a single forward pass taking a fraction of a second, even on a CPU. The proposed model achieves surprisingly good performance on unseen and diverse networks. For example, it is able to predict all 24 million parameters of a ResNet-50 achieving a 60% accuracy on CIFAR-10. On ImageNet, top-5 accuracy of some of our networks approaches 50%. Our task along with the model and results can potentially lead to a new, more computationally efficient paradigm of training networks. Our model also learns a strong representation of neural architectures enabling their analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- From data to functa: Your data point is a function and you can treat it like oneEmilien Dupont, Hyunjik Kim, S. M. Ali Eslami, Danilo Jimenez Rezende 等ICML 2022 · 被引用 209 次
- Equivariant Architectures for Learning in Deep Weight SpacesAviv Navon, Aviv Shamsian, Idan Achituve, Ethan Fetaya 等ICML 2023 · 被引用 101 次
- ATT3D: Amortized Text-to-3D Object SynthesisJonathan Lorraine, Kevin Xie, Xiaohui Zeng, Chen-Hsuan Lin 等ICCV 2023 · 被引用 100 次
- Permutation Equivariant Neural FunctionalsAllan Zhou, Kaien Yang, Kaylee Burns, Adriano Cardace 等NeurIPS 2023 · 被引用 84 次
- Hyper-Representations as Generative Models: Sampling Unseen Neural Network WeightsKonstantin Schürholt, Boris Knyazev, Xavier Giró-i-Nieto, Damian BorthNeurIPS 2022 · 被引用 78 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Geom-GCN: Geometric Graph Convolutional NetworksHongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei 等ICLR 2020 · 被引用 1,445 次
相关 Paper
- A Semi-Supervised Assessor of Neural ArchitecturesYehui Tang, Yunhe Wang, Yixing Xu, Hanting Chen 等CVPR 2020
- Graph Neural Networks for Learning Equivariant Representations of Neural NetworksMiltiadis Kofinas, Boris Knyazev, Yan Zhang, Yunlu Chen 等ICLR 2024 · 被引用 57 次
- Can We Scale Transformers to Predict Parameters of Diverse ImageNet Models?Boris Knyazev, Doha Hwang, Simon Lacoste-JulienICML 2023 · 被引用 31 次
- HyperFast: Instant Classification for Tabular DataDavid Bonet, Daniel Mas Montserrat, Xavier Giró-i-Nieto, Alexander G. IoannidisAAAI 2024 · 被引用 29 次
- Designing Network Design SpacesIlija Radosavovic, Raj Prateek Kosaraju, Ross B. Girshick, Kaiming He 等CVPR 2020
