3D UX-Net: A Large Kernel Volumetric ConvNet Modernizing Hierarchical Transformer for Medical Image Segmentation
Ho Hin Lee, Shunxing Bao, Yuankai Huo, Bennett A. Landman
Abstract
The recent 3D medical ViTs (e.g., SwinUNETR) achieve the state-of-the-art performances on several 3D volumetric data benchmarks, including 3D medical image segmentation. Hierarchical transformers (e.g., Swin Transformers) reintroduced several ConvNet priors and further enhanced the practical viability of adapting volumetric segmentation in 3D medical datasets. The effectiveness of hybrid approaches is largely credited to the large receptive field for non-local selfattention and the large number of model parameters. We hypothesize that volumetric ConvNets can simulate the large receptive field behavior of these learning approaches with fewer model parameters using depth-wise convolution. In this work, we propose a lightweight volumetric ConvNet, termed 3D UX-Net, which adapts the hierarchical transformer using ConvNet modules for robust volumetric segmentation. Specifically, we revisit volumetric depth-wise convolutions with large kernel (LK) size (e.g. starting from 7 × 7 × 7) to enable the larger global receptive fields, inspired by Swin Transformer. We further substitute the multi-layer perceptron (MLP) in Swin Transformer blocks with pointwise depth convolutions and enhance model performances with fewer normalization and activation layers, thus reducing the number of model parameters. 3D UX-Net competes favorably with current SOTA transformers (e.g. SwinUNETR) using three challenging public datasets on volumetric brain and abdominal imaging: 1) MICCAI Challenge 2021 FLARE, 2) MICCAI Challenge 2021 FeTA, and 3) MICCAI Challenge 2022 AMOS. 3D UX-Net consistently outperforms Swin-UNETR with improvement from 0.929 to 0.938 Dice (FLARE2021) and 0.867 to 0.874 Dice (Feta2021). We further evaluate the transfer learning capability of 3D UX-Net with AMOS2022 and demonstrates another improvement of 2.27% Dice (from 0.880 to 0.900). The source code with our proposed model are available at https://github.com/MASILab/3DUX-Net .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 664fc210-66bf-41db-ba1e-5a10f2235782Cited by top-tier papers18
- SegVol: Universal and Interactive Volumetric Medical Image SegmentationYuxin Du, Fan Bai, Tiejun Huang, Bo ZhaoNeurIPS 2024 · 155 citations
- A Robust Mutual-Reinforcing Framework for 3D Multi-Modal Medical Image Fusion Based on Visual-Semantic ConsistencyHao Zhang, Xuhui Zuo, Huabing Zhou, Tao Lu et al.AAAI 2024 · 20 citations
- Towards a Comprehensive, Efficient and Promptable Anatomic Structure Segmentation Model Using 3D Whole-Body CT ScansHeng Guo, Jianfeng Zhang, Jiaxing Huang, Tony C. W. Mok et al.AAAI 2025 · 12 citations
- Evidential Uncertainty-Guided Mitochondria Segmentation for 3D EM ImagesRuohua Shi, Lingyu Duan, Tiejun Huang, Tingting JiangAAAI 2024 · 9 citations
- Upping the Game: How 2D U-Net Skip Connections Flip 3D SegmentationXingru Huang, Yihao Guo, Jian Huang, Tianyun Zhang et al.NeurIPS 2024 · 8 citations
Builds on9
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image AnalysisYucheng Tang, Dong Yang, Wenqi Li, Holger R. Roth et al.CVPR 2022 · 736 citations
- Nested Hierarchical Transformer: Towards Accurate, Data-Efficient and Interpretable Visual UnderstandingZizhao Zhang, Han Zhang, Long Zhao, Ting Chen et al.AAAI 2022 · 216 citations
Related papers
- EffiDec3D: An Optimized Decoder for High-Performance and Efficient 3D Medical Image SegmentationMd Mostafijur Rahman, Radu MarculescuCVPR 2025
- Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentationFenghe Tang, Bingkun Nian, Jianrui Ding, Wenxin Ma et al.ACM MM 2025 · 29 citations
- E2ENet: Dynamic Sparse Feature Fusion for Accurate and Efficient 3D Medical Image SegmentationBoqian Wu, Qiao Xiao, Shiwei Liu, Lu Yin et al.NeurIPS 2024 · 29 citations
- SimpleClick: Interactive Image Segmentation with Simple Vision TransformersQin Liu, Zhenlin Xu, Gedas Bertasius, Marc NiethammerICCV 2023 · 161 citations
- nnWNet: Rethinking the Use of Transformers in Biomedical Image Segmentation and Calling for a Unified Evaluation BenchmarkYanfeng Zhou, Lingrui Li, Le Lu, Minfeng XuCVPR 2025
