EMCAD: Efficient Multi-Scale Convolutional Attention Decoding for Medical Image Segmentation
Md Mostafijur Rahman, Mustafa Munir, Radu Marculescu
Abstract
An efficient and effective decoding mechanism is crucial in medical image segmentation, especially in scenarios with limited computational resources. However, these decoding mechanisms usually come with high computational costs. To address this concern, we introduce EMCAD, a new efficient multi-scale convolutional attention decoder, designed to optimize both performance and computational efficiency. EMCAD leverages a unique multi-scale depth-wise convolution block, significantly enhancing feature maps through multi-scale convolutions. EMCAD also employs channel, spatial, and grouped (large-kernel) gated attention mechanisms, which are highly effective at capturing intricate spatial relationships while focusing on salient regions. By employing group and depth-wise convolution, EMCAD is very efficient and scales well (e.g., only 1.91M parameters and 0.381G FLOPs are needed when using a standard encoder). Our rigorous evaluations across 12 datasets that belong to six medical image segmentation tasks reveal that EMCAD achieves state-of-the-art (SOTA) performance with 79.4% and 80.3% reduction in #Params and #FLOPs, respectively. Moreover, EMCAD's adaptability to different encoders and versatility across segmentation tasks further establish EMCAD as a promising tool, advancing the field towards more efficient and accurate medical image analysis. Our implementation is available at https://github.com/SLDGroupIEMCAD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3590feaf-217e-45bf-a0dc-b02403de6422Cited by top-tier papers18
- Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentationFenghe Tang, Bingkun Nian, Jianrui Ding, Wenxin Ma et al.ACM MM 2025 · 29 citations
- EEO-TFV: Escape-Explore Optimizer for Web-Scale Time-Series Forecasting and Vision AnalysisHua Wang, Jinghao Lu, Fan ZhangWWW 2026 · 6 citations
- Decoding with Structured Awareness: Integrating Directional, Frequency-Spatial, and Structural Attention for Medical Image SegmentationFan Zhang, Zhiwei Gu, Hua WangAAAI 2026 · 4 citations
- LoMix: Learnable Weighted Multi-Scale Logits Mixing for Medical Image SegmentationMd Mostafijur Rahman, Radu MarculescuNeurIPS 2025 · 2 citations
- Breaking Grid Constraints: Dynamic Graph Reconstruction Network for Multi-Organ SegmentationJunhao Xiao, Yang Wei, Jingyu Wang, Yongchao Wang et al.ICCV 2025 · 1 citation
Builds on10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- EffiDec3D: An Optimized Decoder for High-Performance and Efficient 3D Medical Image SegmentationMd Mostafijur Rahman, Radu MarculescuCVPR 2025
- Modality-Agnostic Domain Generalizable Medical Image Segmentation by Multi-Frequency in Multi-Scale AttentionJu-Hyeon Nam, Nur Suriza Syazwany, Su Jung Kim, Sang-Chul LeeCVPR 2024 · 73 citations
- E2ENet: Dynamic Sparse Feature Fusion for Accurate and Efficient 3D Medical Image SegmentationBoqian Wu, Qiao Xiao, Shiwei Liu, Lu Yin et al.NeurIPS 2024 · 29 citations
- S2M-Net: Spectral-Spatial Mixing with Morphology-Aware Adaptive Loss for Medical Image SegmentationSanaullah Chowdhury, Lameya SabrinICML 2026
- ECA-Net: Efficient Channel Attention for Deep Convolutional Neural NetworksQilong Wang, Banggu Wu, Pengfei Zhu, Peihua Li et al.CVPR 2020
