ICML2026
S2M-Net: Spectral-Spatial Mixing with Morphology-Aware Adaptive Loss for Medical Image Segmentation
Sanaullah Chowdhury, Lameya Sabrin
Abstract
Medical image segmentation requires balancing global context with computational efficiency, where self-attention mechanisms suffer from quadratic complexity. We propose S2M-Net, a parameter-efficient architecture (4.7M parameters) that achieves computational savings through Spectral--Spatial Token Mixing (SSTM). SSTM achieves complexity through efficient combination of frequency-domain processing and bottlenecked spatial gating (), exploiting spectral concentration where of energy is captured by low-frequency components (0.8% of the spectrum at resolution). This design avoids self-attention's prohibitive attention map computations while preserving global receptive fields. To handle geometric diversity, we introduce Morphology-Aware Adaptive Segmentation Loss (MASL), which automatically modulates five loss objectives based on per-sample morphological descriptors (tubularity, compactness, irregularity, and scale). Evaluation across 15 datasets spanning 8 modalities demonstrates competitive performance, obtaining the best performance on 14 of 15 datasets, with statistically significant improvements (, Bonferroni-corrected) on 7 challenging tasks (complex morphology, class imbalance, and multi-class segmentation), and clinically meaningful gains (-- Dice) on 8 mature benchmarks. Notably, S2M-Net achieves Dice on EndoVis17 multiclass instrument segmentation ( over TransUNet and over the best baseline UMamba at ), while using fewer parameters (4.7M vs. 60M).