Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentation
Fenghe Tang, Bingkun Nian, Jianrui Ding, Wenxin Ma, Quan Quan, Chengqi Dong, Jie Yang, Wei Liu, S. Kevin Zhou
摘要
In clinical practice, medical image analysis often requires efficient execution on resource-constrained mobile devices. However, existing mobile models-primarily optimized for natural images-tend to perform poorly on medical tasks due to the significant information density gap between natural and medical domains. Combining computational efficiency with medical imaging-specific architectural advantages remains a challenge when developing lightweight, universal, and high-performing networks. To address this, we propose a mobile model called Mobile U-shaped Vision Transformer (Mobile U-ViT) tailored for medical image segmentation. Specifically, we employ the newly proposed ConvUtr as a hierarchical patch embedding, featuring a parameter-efficient large-kernel CNN with inverted bottleneck fusion. This design exhibits transformer-like representation learning capacity while being lighter and faster. To enable efficient local-global information exchange, we introduce a novel Large-kernel Local-Global-Local (LKLGL) block that effectively balances the low information density and high-level semantic discrepancy of medical images. Finally, we incorporate a shallow and lightweight transformer bottleneck for long-range modeling and employ a cascaded decoder with downsampled skip connections for dense prediction. Despite its reduced computational demands, our medical-optimized architecture achieves state-of-the-art performance across eight public 2D and 3D datasets covering diverse imaging modalities, including zero-shot testing on four unseen datasets. These results establish it as an efficient yet powerful and generalization solution for mobile medical image analysis. Code is available at: https://github.com/FengheTan9/Mobile-U-ViT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 被引用 2,162 次
- UCTransNet: Rethinking the Skip Connections in U-Net from a Channel-Wise Perspective with TransformerHaonan Wang, Peng Cao, Jiaqi Wang, Osmar R. ZaïaneAAAI 2022 · 被引用 1,144 次
- Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image AnalysisYucheng Tang, Dong Yang, Wenqi Li, Holger R. Roth 等CVPR 2022 · 被引用 736 次
相关 Paper
- 3D UX-Net: A Large Kernel Volumetric ConvNet Modernizing Hierarchical Transformer for Medical Image SegmentationHo Hin Lee, Shunxing Bao, Yuankai Huo, Bennett A. LandmanICLR 2023 · 被引用 100 次
- Lite Vision Transformer with Enhanced Self-AttentionChenglin Yang, Yilin Wang, Jianming Zhang, He Zhang 等CVPR 2022 · 被引用 139 次
- DeNAS-ViT: Data Efficient NAS-Optimized Vision Transformer for Ultrasound Image SegmentationRenqi Chen, Xinzhe Zheng, Haoyang Su, Kehan WuAAAI 2026 · 被引用 3 次
- Keep It Frozen: Domain-Routed Conditional Residual Modulation for Multi-Domain Vision TransformersUfaq Khan, Umair Nawaz, Massimo Caputo, Muhammad Bilal 等CVPR 2026
- Masked LoGoNet: Fast and Accurate 3D Image Analysis for Medical DomainAmin Karimi Monsefi, Payam Karisani, Mengxi Zhou, Stacey Choi 等KDD 2024 · 被引用 2 次
