Weight-Sharing NAS with Architecture-Agnostic Intermediate Representation
Sixing Yu, Arya Mazaheri, Ali Jannesari
摘要
Weight-sharing supernet has been widely adopted in Neural Architecture Search (NAS) as a promising strategy to obtain smaller and more efficient high-performance models. However, constructing supernets requires domain expertise to design architecture-specific rules (e.g., rules for CNNs and Transformers) for generating subnets, and training a supernet demands joint optimization over a vast sample space of subnets, which is computationally expensive. This paper presents OSF (Optimized Supernet Formation), an automated and architecture-agnostic approach that transforms predefined/pretrained models into weight-sharing supernets. Specifically, we propose representing neural architectures using a high-level computational graph intermediate representation (IR) that enables both the conversion of different types of models into supernets and the extraction of executable subnets via graph traversal. To improve supernet training efficiency, we introduce a sampling strategy that prioritizes the most promising subnet candidates during training, and propose a fork-join parallel training approach with gradient accumulation that resolves write-after-write dependencies, enabling concurrent training of multiple subnet architectures with shared weights. Our empirical evaluations demonstrate that OSF successfully builds supernets from various architectures (CNNs, Transformers, SSMs, and MLPs) while achieving superior performance across language and vision benchmarks. Notably, for Vision Transformers (ViT), OSF reduces FLOPs by 49% while maintaining the accuracy, resulting in a 155% increase in throughput and 35% latency reduction. Code Open-sourced at: https://github.com/yusx-swapp/OSF
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu 等ICCV 2021 · 被引用 2,462 次
相关 Paper
- ShiftNAS: Improving One-shot NAS via Probability ShiftMingyang Zhang, Xinyi Yu, Haodong Zhao, Linlin OuICCV 2023 · 被引用 9 次
- ElasticViT: Conflict-aware Supernet Training for Deploying Fast Vision Transformer on Diverse Mobile DevicesChen Tang, Li Lyna Zhang, Huiqiang Jiang, Jiahang Xu 等ICCV 2023 · 被引用 15 次
- NASViT: Neural Architecture Search for Efficient Vision Transformers with Gradient Conflict aware Supernet TrainingChengyue Gong, Dilin Wang, Meng Li, Xinlei Chen 等ICLR 2022 · 被引用 114 次
- SuperFast: Fast Supernet Training Using Initial KnowledgeMoritz Thoma, Emad Aghajanzadeh, Shambhavi Balamuthu Sampath, Pierpaolo Morì 等DAC 2025
- Rethinking Vision Transformers for MobileNet Size and SpeedYanyu Li, Ju Hu, Yang Wen, Georgios Evangelidis 等ICCV 2023 · 被引用 300 次
