CATE: Computation-aware Neural Architecture Encoding with Transformers
Shen Yan, Kaiqiang Song, Fei Liu, Mi Zhang
摘要
Recent works (White et al., 2020a; Yan et al., 2020) demonstrate the importance of architecture encodings in Neural Architecture Search (NAS). These encodings encode either structure or computation information of the neural architectures. Compared to structure-aware encodings, computation-aware encodings map architectures with similar accuracies to the same region, which improves the downstream architecture search performance (Zhang et al., 2019; White et al., 2020a). In this work, we introduce a Computation-Aware Transformer-based Encoding method called CATE. Different from existing computation-aware encodings based on fixed transformation (e.g. path encoding), CATE employs a pairwise pre-training scheme to learn computation-aware encodings using Transformers with cross-attention. Such learned encodings contain dense and contextualized computation information of neural architectures. We compare CATE with eleven encodings under three major encoding-dependent NAS subroutines in both small and large search spaces. Our experiments show that CATE is beneficial to the downstream search, especially in the large search space. Moreover, the outside search space experiment demonstrates its superior generalization ability beyond the search space on which it was trained. Our code is available at: https://github.com/MSU-MLSys-Lab/CATE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- NAS-Bench-x11 and the Power of Learning CurvesShen Yan, Colin White, Yash Savani, Frank HutterNeurIPS 2021 · 被引用 36 次
- PINAT: A Permutation INvariance Augmented Transformer for NAS PredictorShun Lu, Yu Hu, Peihao Wang, Yan Han 等AAAI 2023 · 被引用 31 次
- Efficient Data Subset Selection to Generalize Training Across Models: Transductive and Inductive NetworksEeshaan Jain, Tushar Nandy, Gaurav Aggarwal, Ashish Tendulkar 等NeurIPS 2023 · 被引用 30 次
- DiffusionNAG: Predictor-guided Neural Architecture Generation with Diffusion ModelsSohyun An, Hayeon Lee, Jaehyeong Jo, Seanie Lee 等ICLR 2024 · 被引用 21 次
- TA-GATES: An Encoding Scheme for Neural Network ArchitecturesXuefei Ning, Zixuan Zhou, Junbo Zhao, Tianchen Zhao 等NeurIPS 2022 · 被引用 19 次
它引用的顶会 Paper17
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 被引用 825 次
- Understanding and Robustifying Differentiable Architecture SearchArber Zela, Thomas Elsken, Tonmoy Saikia, Yassine Marrakchi 等ICLR 2020 · 被引用 408 次
- BANANAS: Bayesian Optimization with Neural Architectures for Neural Architecture SearchColin White, Willie Neiswanger, Yash SavaniAAAI 2021 · 被引用 401 次
- Exploring Randomly Wired Neural Networks for Image RecognitionSaining Xie, Alexander Kirillov, Ross B. Girshick, Kaiming HeICCV 2019 · 被引用 384 次
- Stabilizing Differentiable Architecture Search via Perturbation-based RegularizationXiangning Chen, Cho-Jui HsiehICML 2020 · 被引用 235 次
相关 Paper
- Does Unsupervised Architecture Representation Learning Help Neural Architecture Search?Shen Yan, Yu Zheng, Wei Ao, Xiao Zeng 等NeurIPS 2020 · 被引用 129 次
- Encodings for Prediction-based Neural Architecture SearchYash Akhauri, Mohamed S. AbdelfattahICML 2024 · 被引用 8 次
- TNASP: A Transformer-based NAS Predictor with a Self-evolution FrameworkShun Lu, Jixiang Li, Jianchao Tan, Sen Yang 等NeurIPS 2021 · 被引用 51 次
- Prior Knowledge Guided Neural Architecture GenerationJingrong Xie, Han Ji, Yanan SunICML 2025
- A Study on Encodings for Neural Architecture SearchColin White, Willie Neiswanger, Sam Nolen, Yash SavaniNeurIPS 2020 · 被引用 88 次
