HR-NAS: Searching Efficient High-Resolution Neural Architectures With Lightweight Transformers
Mingyu Ding, Xiaochen Lian, Linjie Yang, Peng Wang, Xiaojie Jin, Zhiwu Lu, Ping Luo
Abstract
High-resolution representations (HR) are essential for dense prediction tasks such as segmentation, detection, and pose estimation. Learning HR representations is typically ignored in previous Neural Architecture Search (NAS) methods that focus on image classification. This work proposes a novel NAS method, called HR-NAS, which is able to find efficient and accurate networks for different tasks, by effectively encoding multiscale contextual information while maintaining high-resolution representations. In HR-NAS, we renovate the NAS search space as well as its searching strategy. To better encode multiscale image contexts in the search space of HR-NAS, we first carefully design a lightweight transformer, whose computational complexity can be dynamically changed with respect to different objective functions and computation budgets. To maintain highresolution representations of the learned networks, HR-NAS adopts a multi-branch architecture that provides convolutional encoding of multiple feature resolutions, inspired by HRNet [76]. Last, we proposed an efficient fine-grained search strategy to train HR-NAS, which effectively explores the search space, and finds optimal architectures given various tasks and computation resources. As shown in Fig. 1 (a), HR-NAS is capable of achieving state-of-the-art tradeoffs between performance and FLOPs for three dense prediction tasks and an image classification task, given only small computational budgets. For example, HR-NAS surpasses SqueezeNAS [66] that is specially designed for semantic segmentation while improving efficiency by 45.9%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79a59857-be52-4e82-b093-a7813f1e3c87Cited by top-tier papers14
- TopFormer: Token Pyramid Transformer for Mobile Semantic SegmentationWenqiang Zhang, Zilong Huang, Guozhong Luo, Tao Chen et al.CVPR 2022 · 313 citations
- PhysFormer: Facial Video-based Physiological Measurement with Temporal Difference TransformerZitong Yu, Yuming Shen, Jingang Shi, Hengshuang Zhao et al.CVPR 2022 · 255 citations
- Multi-Scale High-Resolution Vision Transformer for Semantic SegmentationJiaqi Gu, Hyoukjun Kwon, Dilin Wang, Wei Ye et al.CVPR 2022 · 236 citations
- Reversible Column NetworksYuxuan Cai, Yizhuang Zhou, Qi Han, Jianjian Sun et al.ICLR 2023 · 21 citations
- Towards Real-Time Segmentation on the EdgeYanyu Li, Changdi Yang, Pu Zhao, Geng Yuan et al.AAAI 2023 · 19 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
Related papers
- HRFormer: High-Resolution Vision Transformer for Dense PredictYuhui Yuan, Rao Fu, Lang Huang, Weihong Lin et al.NeurIPS 2021 · 357 citations
- Pose-native Network Architecture Search for Multi-person Human Pose EstimationQian Bao, Wu Liu, Jun Hong, Lingyu Duan et al.ACM MM 2020 · 13 citations
- DCNAS: Densely Connected Neural Architecture Search for Semantic Image SegmentationXiong Zhang, Hongmin Xu, Hong Mo, Jianchao Tan et al.CVPR 2021
- Exploring Relational Context for Multi-Task Dense PredictionDavid Brüggemann, Menelaos Kanakis, Anton Obukhov, Stamatios Georgoulis et al.ICCV 2021 · 110 citations
- ViPNAS: Efficient Video Pose Estimation via Neural Architecture SearchLumin Xu, Yingda Guan, Sheng Jin, Wentao Liu et al.CVPR 2021
