EffiDec3D: An Optimized Decoder for High-Performance and Efficient 3D Medical Image Segmentation
Md Mostafijur Rahman, Radu Marculescu
Abstract
Recent 3D deep networks such as SwinUNETR, Swin-UNETRv2, and 3D UX-Net have shown promising performance by leveraging self-attention and large-kernel convolutions to capture the volumetric context. However, their substantial computational requirements limit their use in real-time and resource-constrained environments. The high #FLOPs and #Params in these networks stem largely from complex decoder designs with high-resolution layers and excessive channel counts. In this paper, we propose Ef-fiDec3D, an optimized 3D decoder that employs a channel reduction strategy across all decoder stages, which sets the number of channels to the minimum needed for accurate feature representation. Additionally, EffiDec3D removes the high-resolution layers when their contribution to segmentation quality is minimal. Our optimized Ef-fiDec3D decoder achieves a 96.4% reduction in #Params and a 93.0% reduction in #FLOPs compared to the decoder of original 3D UX-Net. Similarly, for SwinUNETR and SwinUNETRv2 (which share an identical decoder), we observe reductions of 94.9% in #Params and 86.2% in #FLOPs. Our extensive experiments on 12 different medical imaging tasks confirm that EffiDec3D not only significantly reduces the computational demands, but also maintains a performance level comparable to original models, thus establishing a new standard for efficient 3D medical image segmentation. Our implementation is available at https://github.com/SLDGroup/EffiDec3D.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image AnalysisYucheng Tang, Dong Yang, Wenqi Li, Holger R. Roth et al.CVPR 2022 · 736 citations
- EMCAD: Efficient Multi-Scale Convolutional Attention Decoding for Medical Image SegmentationMd Mostafijur Rahman, Mustafa Munir, Radu MarculescuCVPR 2024 · 352 citations
- 3D UX-Net: A Large Kernel Volumetric ConvNet Modernizing Hierarchical Transformer for Medical Image SegmentationHo Hin Lee, Shunxing Bao, Yuankai Huo, Bennett A. LandmanICLR 2023 · 100 citations
Related papers
- E2ENet: Dynamic Sparse Feature Fusion for Accurate and Efficient 3D Medical Image SegmentationBoqian Wu, Qiao Xiao, Shiwei Liu, Lu Yin et al.NeurIPS 2024 · 29 citations
- Upping the Game: How 2D U-Net Skip Connections Flip 3D SegmentationXingru Huang, Yihao Guo, Jian Huang, Tianyun Zhang et al.NeurIPS 2024 · 8 citations
- Non-Local U-Nets for Biomedical Image SegmentationZhengyang Wang, Na Zou, Dinggang Shen, Shuiwang JiAAAI 2020 · 180 citations
- Efficient Folded Attention for Medical Image Reconstruction and SegmentationHang Zhang, Jinwei Zhang, Rongguang Wang, Qihao Zhang et al.AAAI 2021 · 23 citations
- ShiftMorph: A Fast and Robust Convolutional Neural Network for 3D Deformable Medical Image RegistrationLijian Yang, Weisheng Li, Yucheng Shu, Jian-Xun Mi et al.ACM MM 2024 · 4 citations
