ASMIL: Attention-Stabilized Multiple Instance Learning for Whole-Slide Imaging
Linfeng Ye, Shayan Mohajer Hamidi, Zhixiang Chi, Guang Li, Mert Pilanci, Takahiro Ogawa, Miki Haseyama, Konstantinos N. Plataniotis
Abstract
Attention-based multiple instance learning (MIL) has emerged as a powerful framework for whole slide image (WSI) diagnosis, leveraging attention to aggregate instance-level features into bag-level predictions. Despite this success, we find that such methods exhibit a new failure mode: unstable attention dynamics. Across four representative attention-based MIL methods and two public WSI datasets, we observe that attention distributions oscillate across epochs rather than converging to a consistent pattern, degrading performance. This instability adds to two previously reported challenges: overfitting and over-concentrated attention distribution. To simultaneously overcome these three limitations, we introduce attention-stabilized multiple instance learning (ASMIL), a novel unified framework. ASMIL uses an anchor model to stabilize attention, replaces softmax with a normalized sigmoid function in the anchor to prevent over-concentration, and applies token random dropping to mitigate overfitting. Extensive experiments demonstrate that ASMIL achieves up to a 6.49% F1 score improvement over state-of-the-art methods. Moreover, integrating the anchor model and normalized sigmoid into existing attention-based MIL methods consistently boosts their performance, with F1 score gains up to 10.73%. All code and data are publicly available at https://anonymous.4open.science/r/ASMIL-5018/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa075f65-fae2-47fc-8a7a-a287bc8417f7Cited by top-tier papers2
- Widget2Code: From Visual Widgets to UI Code via Multimodal LLMsHouston H. Zhang, Tao Zhang, Baoze Lin, Yuanqi Xue et al.CVPR 2026 · 9 citations
- Bridging the Modality Gap in Compositional Zero-Shot Learning via Sparse Alignment and Unimodal Memory BankYang Zhang, Zhixiang Chi, Xudong Yan, Yang Wang et al.CVPR 2026
Builds on27
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image ClassificationZhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang et al.NeurIPS 2021 · 1,163 citations
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
Related papers
- How Effective Can Dropout Be in Multiple Instance Learning ?Wenhui Zhu, Peijie Qiu, Xiwen Chen, Zhangsihao Yang et al.ICML 2025
- Bayes-MIL: A New Probabilistic Perspective on Attention-based Multiple Instance Learning for Whole Slide ImagesYufei Cui, Ziquan Liu, Xiangyu Liu, Xue Liu et al.ICLR 2023
- Bi-directional Weakly Supervised Knowledge Distillation for Whole Slide Image ClassificationLinhao Qu, Xiaoyuan Luo, Manning Wang, Zhijian SongNeurIPS 2022 · 88 citations
- Boosting Multiple Instance Learning Models for Whole Slide Image Classification: A Model-Agnostic Framework Based on Counterfactual InferenceWeiping Lin, Zhenfeng Zhuang, Lequan Yu, Liansheng WangAAAI 2024 · 19 citations
- Dual-Curriculum Contrastive Multi-Instance Learning for Cancer Prognosis Analysis with Whole Slide ImagesChao Tu, Yu Zhang, Zhenyuan NingNeurIPS 2022 · 24 citations
