Learning to Flow from Generative Pretext Tasks for Neural Architecture Encoding
Sunwoo Kim, Hyunjin Hwang, Kijung Shin
摘要
The performance of a deep learning model on a specific task and dataset depends heavily on its neural architecture, motivating considerable efforts to rapidly and accurately identify architectures suited to the target task and dataset. To achieve this, researchers use machine learning models-typically neural architecture encoders-to predict the performance of a neural architecture. Many state-of-the-art encoders aim to capture information flow within a neural architecture, which reflects how information moves through the forward pass and backpropagation, via a specialized model structure. However, due to their complicated structures, these flow-based encoders are significantly slower to process neural architectures compared to simpler encoders, presenting a notable practical challenge. To address this, we propose FGP, a novel pre-training method for neural architecture encoding that trains an encoder to capture the information flow without requiring specialized model structures. FGP trains an encoder to reconstruct a flow surrogate, our proposed representation of the neural architecture's information flow. Our experiments show that FGP boosts encoder performance by up to 106% in Precision@1%, compared to the same encoder trained solely with supervised learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen 等NeurIPS 2020 · 被引用 3,042 次
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik 等ICLR 2020 · 被引用 1,744 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 被引用 825 次
相关 Paper
- Does Unsupervised Architecture Representation Learning Help Neural Architecture Search?Shen Yan, Yu Zheng, Wei Ao, Xiao Zeng 等NeurIPS 2020 · 被引用 129 次
- FlowerFormer: Empowering Neural Architecture Encoding Using a Flow-Aware Graph TransformerDongyeong Hwang, Hyunju Kim, Sunwoo Kim, Kijung ShinCVPR 2024
- AIO-P: Expanding Neural Performance Predictors beyond Image ClassificationKeith G. Mills, Di Niu, Mohammad Salameh, Weichen Qiu 等AAAI 2023 · 被引用 9 次
- A Semi-Supervised Assessor of Neural ArchitecturesYehui Tang, Yunhe Wang, Yixing Xu, Hanting Chen 等CVPR 2020
- FlowGNN: A Dataflow Architecture for Real-Time Workload-Agnostic Graph Neural Network InferenceRishov Sarkar, Stefan Abi-Karam, Yuqi He, Lakshmi Sathidevi 等HPCA 2023 · 被引用 100 次
