LoMix: Learnable Weighted Multi-Scale Logits Mixing for Medical Image Segmentation
Md Mostafijur Rahman, Radu Marculescu
摘要
U-shaped networks output logits at multiple spatial scales, each capturing a different blend of coarse context and fine detail. Yet, training still treats these logits in isolation-either supervising only the final, highest-resolution logits or applying deep supervision with identical loss weights at every scale-without exploring mixed-scale combinations. Consequently, the decoder output misses the complementary cues that arise only when coarse and fine predictions are fused. To address this issue, we introduce LoMix (Logits Mixing), a Neural Architecture Search (NAS)-inspired, differentiable plug-and-play module that generates new mixed-scale outputs and learns how exactly each of them should guide the training process. More precisely, LoMix mixes the multi-scale decoder logits with four lightweight fusion operators: addition, multiplication, concatenation, and attentionbased weighted fusion, yielding a rich set of synthetic "mutant" maps. Every original or mutant map is given a softplus loss weight that is co-optimized with network parameters, mimicking a one-step architecture search that automatically discovers the most useful scales, mixtures, and operators. Plugging LoMix into recent U-shaped architectures (i.e., PVT-V2-B2 backbone with EMCAD decoder) on Synapse 8-organ dataset improves DICE by +4.2% over single-output supervision, +2.2% over deep supervision, and +1.5% over equally weighted additive fusion, all with zero inference overhead. When training data are scarce (e.g., one or two labeled scans, 5% of the trainset), the advantage grows to +9.23%, underscoring LoMix's data efficiency. Across four benchmarks and diverse U-shaped networks, LoMiX improves DICE by up to +13.5% over single-output supervision, confirming that learnable weighted mixed-scale fusion generalizes broadly while remaining data efficient, fully interpretable, and overhead-free at inference. Our implementation is available at https://github.com/SLDGroup/LoMix.
2 Related Work Medical Segmentation Architectures: U-shaped encoder-decoder networks with skip connections (e.g., U-Net) are the de facto architectures in medical image segmentation [25]; they can capture fine details via multi-scale feature maps, but are limited by the locality of convolutional operations. For example, Chen et al. note that standard U-Net struggles to model long-range dependencies, motivating hybrid designs [3]. To address this, recent work has proposed transformer-based backbones for segmentation. TransUNet [3] combines a CNN encoder with a Vision Transformer to learn the global context, while the decoder recovers the spatial details. Similarly, Swin-Unet [2] uses a hierarchical Swin Transformer in both encoder and decoder, demonstrating that pure-transformer U-shaped models outperform purely convolutional ones on multi-organ tasks.
Other variants adopt the Pyramid Vision Transformer (PVT) [34] as the encoder. For instance, CASCADE use PVT encoders with novel attention-or graph-based decoders to progressively refine multi-scale features. Polyp-PVT [7] and SSFormer [33] also leverage PVT backbones for polyp segmentation, incorporating hand-crafted fusion modules (e.g., cascaded fusion, camouflage, and locality decoders) to combine features across scales.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- EMCAD: Efficient Multi-Scale Convolutional Attention Decoding for Medical Image SegmentationMd Mostafijur Rahman, Mustafa Munir, Radu MarculescuCVPR 2024 · 被引用 352 次
- UACANet: Uncertainty Augmented Context Attention for Polyp SegmentationTaehun Kim, Hyemin Lee, Daijin KimACM MM 2021 · 被引用 305 次
- Rolling-Unet: Revitalizing MLP's Ability to Efficiently Extract Long-Distance Dependencies for Medical Image SegmentationYutong Liu, Haijiang Zhu, Mengting Liu, Huaiyuan Yu 等AAAI 2024 · 被引用 136 次
相关 Paper
- UCTransNet: Rethinking the Skip Connections in U-Net from a Channel-Wise Perspective with TransformerHaonan Wang, Peng Cao, Jiaqi Wang, Osmar R. ZaïaneAAAI 2022 · 被引用 1,144 次
- Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentationFenghe Tang, Bingkun Nian, Jianrui Ding, Wenxin Ma 等ACM MM 2025 · 被引用 29 次
- Rethinking U-Net: Task-Adaptive Mixture of Skip Connections for Enhanced Medical Image SegmentationZichen Luo, Xinshan Zhu, Lan Zhang, Biao SunAAAI 2025 · 被引用 8 次
- DeNAS-ViT: Data Efficient NAS-Optimized Vision Transformer for Ultrasound Image SegmentationRenqi Chen, Xinzhe Zheng, Haoyang Su, Kehan WuAAAI 2026 · 被引用 3 次
- UniFuse: A Unified All-In-One Framework for Multi-Modal Medical Image Fusion Under Diverse Degradations and MisalignmentsDayong Su, Yafei Zhang, Huafeng Li, Jinxing Li 等ICCV 2025 · 被引用 2 次
