LoMix: Learnable Weighted Multi-Scale Logits Mixing for Medical Image Segmentation
Md Mostafijur Rahman, Radu Marculescu
Abstract
U-shaped networks output logits at multiple spatial scales, each capturing a different blend of coarse context and fine detail. Yet, training still treats these logits in isolation-either supervising only the final, highest-resolution logits or applying deep supervision with identical loss weights at every scale-without exploring mixed-scale combinations. Consequently, the decoder output misses the complementary cues that arise only when coarse and fine predictions are fused. To address this issue, we introduce LoMix (Logits Mixing), a Neural Architecture Search (NAS)-inspired, differentiable plug-and-play module that generates new mixed-scale outputs and learns how exactly each of them should guide the training process. More precisely, LoMix mixes the multi-scale decoder logits with four lightweight fusion operators: addition, multiplication, concatenation, and attentionbased weighted fusion, yielding a rich set of synthetic "mutant" maps. Every original or mutant map is given a softplus loss weight that is co-optimized with network parameters, mimicking a one-step architecture search that automatically discovers the most useful scales, mixtures, and operators. Plugging LoMix into recent U-shaped architectures (i.e., PVT-V2-B2 backbone with EMCAD decoder) on Synapse 8-organ dataset improves DICE by +4.2% over single-output supervision, +2.2% over deep supervision, and +1.5% over equally weighted additive fusion, all with zero inference overhead. When training data are scarce (e.g., one or two labeled scans, 5% of the trainset), the advantage grows to +9.23%, underscoring LoMix's data efficiency. Across four benchmarks and diverse U-shaped networks, LoMiX improves DICE by up to +13.5% over single-output supervision, confirming that learnable weighted mixed-scale fusion generalizes broadly while remaining data efficient, fully interpretable, and overhead-free at inference. Our implementation is available at https://github.com/SLDGroup/LoMix.
2 Related Work Medical Segmentation Architectures: U-shaped encoder-decoder networks with skip connections (e.g., U-Net) are the de facto architectures in medical image segmentation [25]; they can capture fine details via multi-scale feature maps, but are limited by the locality of convolutional operations. For example, Chen et al. note that standard U-Net struggles to model long-range dependencies, motivating hybrid designs [3]. To address this, recent work has proposed transformer-based backbones for segmentation. TransUNet [3] combines a CNN encoder with a Vision Transformer to learn the global context, while the decoder recovers the spatial details. Similarly, Swin-Unet [2] uses a hierarchical Swin Transformer in both encoder and decoder, demonstrating that pure-transformer U-shaped models outperform purely convolutional ones on multi-organ tasks.
Other variants adopt the Pyramid Vision Transformer (PVT) [34] as the encoder. For instance, CASCADE use PVT encoders with novel attention-or graph-based decoders to progressively refine multi-scale features. Polyp-PVT [7] and SSFormer [33] also leverage PVT backbones for polyp segmentation, incorporating hand-crafted fusion modules (e.g., cascaded fusion, camouflage, and locality decoders) to combine features across scales.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60cd6cfd-4435-4d39-af23-bc5ba5894743Builds on3
- EMCAD: Efficient Multi-Scale Convolutional Attention Decoding for Medical Image SegmentationMd Mostafijur Rahman, Mustafa Munir, Radu MarculescuCVPR 2024 · 352 citations
- UACANet: Uncertainty Augmented Context Attention for Polyp SegmentationTaehun Kim, Hyemin Lee, Daijin KimACM MM 2021 · 305 citations
- Rolling-Unet: Revitalizing MLP's Ability to Efficiently Extract Long-Distance Dependencies for Medical Image SegmentationYutong Liu, Haijiang Zhu, Mengting Liu, Huaiyuan Yu et al.AAAI 2024 · 136 citations
Related papers
- UCTransNet: Rethinking the Skip Connections in U-Net from a Channel-Wise Perspective with TransformerHaonan Wang, Peng Cao, Jiaqi Wang, Osmar R. ZaïaneAAAI 2022 · 1,144 citations
- Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentationFenghe Tang, Bingkun Nian, Jianrui Ding, Wenxin Ma et al.ACM MM 2025 · 29 citations
- Rethinking U-Net: Task-Adaptive Mixture of Skip Connections for Enhanced Medical Image SegmentationZichen Luo, Xinshan Zhu, Lan Zhang, Biao SunAAAI 2025 · 8 citations
- DeNAS-ViT: Data Efficient NAS-Optimized Vision Transformer for Ultrasound Image SegmentationRenqi Chen, Xinzhe Zheng, Haoyang Su, Kehan WuAAAI 2026 · 3 citations
- UniFuse: A Unified All-In-One Framework for Multi-Modal Medical Image Fusion Under Diverse Degradations and MisalignmentsDayong Su, Yafei Zhang, Huafeng Li, Jinxing Li et al.ICCV 2025 · 2 citations
