DeNAS-ViT: Data Efficient NAS-Optimized Vision Transformer for Ultrasound Image Segmentation
Renqi Chen, Xinzhe Zheng, Haoyang Su, Kehan Wu
Abstract
Accurate segmentation of ultrasound images is essential for reliable medical diagnoses but is challenged by poor image quality and scarce labeled data. Prior approaches have relied on manually designed, complex network architectures to improve multi-scale feature extraction. However, such handcrafted models offer limited gains when prior knowledge is inadequate and are prone to overfitting on small datasets. In this paper, we introduce DeNAS-ViT, a Data efficient NAS-optimized Vision Transformer, the first method to leverage neural architecture search (NAS) for ultrasound image segmentation by automatically optimizing model architecture through token-level search. Specifically, we propose an efficient NAS module that performs multi-scale token search prior to the ViT’s attention mechanism, effectively capturing both contextual and local features while minimizing computational costs. Given ultrasound’s data scarcity and NAS’s inherent data demands, we further develop a NAS-guided semi-supervised learning (SSL) framework. This approach integrates network independence and contrastive learning within a stage-wise optimization strategy, significantly enhancing model robustness under limited-data conditions. Extensive experiments on public datasets demonstrate that DeNAS-ViT achieves state-of-the-art performance, maintaining robustness with minimal labeled data. Moreover, we highlight DeNAS-ViT’s generalization potential beyond ultrasound imaging, underscoring its broader applicability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ac796db-0630-4eab-9c90-591dadcf9166Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture SearchYuhui Xu, Lingxi Xie, Xiaopeng Zhang, Xin Chen et al.ICLR 2020 · 691 citations
- EfficientViT: Lightweight Multi-Scale Attention for High-Resolution Dense PredictionHan Cai, Junyan Li, Muyan Hu, Chuang Gan et al.ICCV 2023 · 265 citations
- Rethinking Semi-Supervised Medical Image Segmentation: A Variance-Reduction PerspectiveChenyu You, Weicheng Dai, Yifei Min, Fenglin Liu et al.NeurIPS 2023 · 147 citations
- CauSSL: Causality-inspired Semi-supervised Learning for Medical Image SegmentationJuzheng Miao, Cheng Chen, Furui Liu, Hao Wei et al.ICCV 2023 · 88 citations
Related papers
- Anatomy-aware Representation Learning for Medical UltrasoundSeok-Hwan Oh, Myeong-Gee Kim, Guil Jung, Hyeon-Jik Lee et al.ICLR 2026 · 23 citations
- NASViT: Neural Architecture Search for Efficient Vision Transformers with Gradient Conflict aware Supernet TrainingChengyue Gong, Dilin Wang, Meng Li, Xinlei Chen et al.ICLR 2022 · 114 citations
- Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentationFenghe Tang, Bingkun Nian, Jianrui Ding, Wenxin Ma et al.ACM MM 2025 · 29 citations
- ElasticViT: Conflict-aware Supernet Training for Deploying Fast Vision Transformer on Diverse Mobile DevicesChen Tang, Li Lyna Zhang, Huiqiang Jiang, Jiahang Xu et al.ICCV 2023 · 15 citations
- GLiT: Neural Architecture Search for Global and Local Image TransformerBoyu Chen, Peixia Li, Chuming Li, Baopu Li et al.ICCV 2021 · 100 citations
