NAR-Former V2: Rethinking Transformer for Universal Neural Network Representation Learning
Yun Yi, Haokui Zhang, Rong Xiao, Nannan Wang, Xiaoyu Wang
摘要
As more deep learning models are being applied in real-world applications, there is a growing need for modeling and learning the representations of neural networks themselves. An efficient representation can be used to predict target attributes of networks without the need for actual training and deployment procedures, facilitating efficient network deployment and design. Recently, inspired by the success of Transformer, some Transformer-based representation learning frameworks have been proposed and achieved promising performance in handling cell-structured models. However, graph neural network (GNN) based approaches still dominate the field of learning representation for the entire network. In this paper, we revisit Transformer and compare it with GNN to analyse their different architecture characteristics. We then propose a modified Transformer-based universal neural network representation learning model NAR-Former V2. It can learn efficient representations from both cell-structured networks and entire networks. Specifically, we first take the network as a graph and design a straightforward tokenizer to encode the network into a sequence. Then, we incorporate the inductive representation learning capability of GNN into Transformer, enabling Transformer to generalize better when encountering unseen architecture. Additionally, we introduce a series of simple yet effective modifications to enhance the ability of the Transformer in learning representation from graph structures. Our proposed method surpasses the GNN-based method NNLP by a significant margin in latency estimation on the NNLQP dataset. Furthermore, regarding accuracy prediction on the NASBench101 and NASBench201 datasets, our method achieves highly comparable performance to other state-of-the-art methods. Code is available at https://github.com/yuny220/NAR-Former-V2 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning to Flow from Generative Pretext Tasks for Neural Architecture EncodingSunwoo Kim, Hyunjin Hwang, Kijung ShinNeurIPS 2025 · 被引用 2 次
- NN-Former: Rethinking Graph Structure in Neural Architecture RepresentationRuihan Xu, Haokui Zhang, Yaowei Wang, Wei Zeng 等CVPR 2025
- Biologically Plausible Brain Graph TransformerCiyuan Peng, Yuelong Huang, Qichao Dong, Shuo Yu 等ICLR 2025
它引用的顶会 Paper15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 被引用 825 次
相关 Paper
- NAR-Former: Neural Architecture Representation Learning Towards Holistic Attributes PredictionYun Yi, Haokui Zhang, Wenze Hu, Nannan Wang 等CVPR 2023
- FlowerFormer: Empowering Neural Architecture Encoding Using a Flow-Aware Graph TransformerDongyeong Hwang, Hyunju Kim, Sunwoo Kim, Kijung ShinCVPR 2024
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng 等NeurIPS 2021 · 被引用 1,632 次
- Primphormer: Efficient Graph Transformers with Primal RepresentationsMingzhen He, Ruikai Yang, Hanling Tian, Youmei Qiu 等ICML 2025
- Structure-Aware Transformer for Graph Representation LearningDexiong Chen, Leslie O'Bray, Karsten M. BorgwardtICML 2022 · 被引用 349 次
