Robust Self-Supervised Multi-Instance Learning with Structure Awareness
Yejiang Wang, Yuhai Zhao, Zhengkui Wang, Meixia Wang
Abstract
Multi-instance learning (MIL) is a supervised learning where each example is a labeled bag with many instances. The typical MIL strategies are to train an instance-level feature extractor followed by aggregating instances features as bag-level representation with labeled information. However, learning such a bag-level representation highly depends on a large number of labeled datasets, which are difficult to get in real-world scenarios. In this paper, we make the first attempt to propose a robust Self-supervised Multi-Instance LEarning architecture with Structure awareness (SMILEs) that learns unsupervised bag representation. Our proposed approach is: 1) permutation invariant to the order of instances in bag; 2) structure-aware to encode the topological structures among the instances; and 3) robust against instances noise or permutation. Specifically, to yield robust MIL model without label information, we augment the multi-instance bag and train the representation encoder to maximize the correspondence between the representations of the same bag in its different augmented forms. Moreover, to capture topological structures from nearby instances in bags, our framework learns optimal graph structures for the bags and these graphs are optimized together with message passing layers and the ordered weighted averaging operator towards contrastive loss. Our main theorem characterizes the permutation invariance of the bag representation. Compared with state-of-the-art supervised MIL baselines, SMILEs achieves average improvement of 4.9%, 4.4% in classification accuracy on 5 benchmark datasets and 20 newsgroups datasets, respectively. In addition, we show that the model is robust to the input corruption.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- GALOPA: Graph Transport Learning with Optimal Plan AlignmentYejiang Wang, Yuhai Zhao, Daniel Zhengkui Wang, Ling LiNeurIPS 2023 · 15 citations
- Limited-Supervised Multi-Label Learning with Dependency NoiseYejiang Wang, Yuhai Zhao, Zhengkui Wang, Wen Shan et al.AAAI 2024 · 7 citations
- Utterance-level Emotion Recognition in Conversation with Conversation-level SupervisionXiming Li, Yuanchao Dai, Zhiyao Yang, Jinjin Chi et al.AAAI 2025 · 2 citations
Builds on8
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie et al.NeurIPS 2020 · 1,113 citations
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel et al.NeurIPS 2021 · 421 citations
- Interventional Multi-Instance Learning with Deconfounded Instance-Level PredictionTiancheng Lin, Hongteng Xu, Canqian Yang, Yi XuAAAI 2022 · 34 citations
- Set Norm and Equivariant Skip Connections: Putting the Deep in Deep SetsLily H. Zhang, Veronica Tozzo, John M. Higgins, Rajesh RanganathICML 2022 · 28 citations
- Bag Graph: Multiple Instance Learning Using Bayesian Graph Neural NetworksSoumyasundar Pal, Antonios Valkanas, Florence Regol, Mark CoatesAAAI 2022 · 24 citations
Related papers
- Multiple-Instance Learning from Similar and Dissimilar BagsLei Feng, Senlin Shu, Yuzhou Cao, Lue Tao et al.KDD 2021 · 11 citations
- Self-Supervised Contrastive Re-Learning for Multi-Graph Multi-Label ClassificationMeixia Wang, Yuhai Zhao, Zhengkui Wang, Yejiang Wang et al.AAAI 2026
- Multi-Instance Causal Representation Learning for Instance Label Prediction and Out-of-Distribution GeneralizationWeijia Zhang, Xuanhui Zhang, Hanwen Deng, Min-Ling ZhangNeurIPS 2022 · 32 citations
- Rethinking Multi-Instance Learning Through Graph-Driven Fusion: A Dual-Path Approach to Adaptive RepresentationYu-Xuan Zhang, Zhengchun Zhou, Weisha Liu, Mingxing ZhangAAAI 2026 · 1 citation
- Predicting Lymph Node Metastasis Using Histopathological Images Based on Multiple Instance Learning With Deep Graph ConvolutionYu Zhao, Fan Yang, Yuqi Fang, Hailing Liu et al.CVPR 2020
