Differentiable Expectation-Maximization for Set Representation Learning
Minyoung Kim
Abstract
We tackle the set2vec problem, the task of extracting a vector representation from an input set comprised of a variable number of feature vectors. Although recent approaches based on self attention such as (Set)Transformers were very successful due to the capability of capturing complex interaction between set elements, the computational overhead is the well-known downside. The inducing-point attention and the latest optimal transport kernel embedding (OTKE) are promising remedies that attain comparable or better performance with reduced computational cost, by incorporating a fixed number of learnable queries in attention. In this paper we approach the set2vec problem from a completely different perspective. The elements of an input set are considered as i.i.d. samples from a mixture distribution, and we define our set embedding feed-forward network as the maximum-a-posterior (MAP) estimate of the mixture which is approximately attained by a few Expectation-Maximization (EM) steps. The whole MAP-EM steps are differentiable operations with a fixed number of mixture parameters, allowing efficient auto-diff back-propagation for any given downstream task. Furthermore, the proposed mixture set data fitting framework allows unsupervised set representation learning naturally via marginal likelihood maximization aka the empirical Bayes. Interestingly, we also find that OTKE can be seen as a special case of our framework, specifically a single-step EM with extra balanced assignment constraints on the E-step. Compared to OTKE, our approach provides more flexible set embedding as well as prior-induced model regularization. We evaluate our approach on various tasks demonstrating improved performance over the state-of-the-arts.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers13
- Multimodal Prototyping for cancer survival predictionAndrew H. Song, Richard J. Chen, Guillaume Jaume, Anurag J. Vaidya et al.ICML 2024 · 53 citations
- Morphological Prototyping for Unsupervised Slide Representation Learning in Computational PathologyAndrew H. Song, Richard J. Chen, Tong Ding, Drew F. K. Williamson et al.CVPR 2024 · 51 citations
- DPsurv: Dual-Prototype Evidential Fusion for Uncertainty-Aware and Interpretable Whole Slide Image Survival PredictionYucheng Xing, ling huang, Jingying Ma, Ruping Hong et al.ICML 2026 · 8 citations
- Fisher Information Embedding for Node and Graph LearningDexiong Chen, Paolo Pellizzoni, Karsten M. BorgwardtICML 2023 · 4 citations
- Scalable Set Encoding with Universal Mini-Batch Consistency and Unbiased Full Set Gradient ApproximationJeffrey Willette, Seanie Lee, Bruno Andreis, Kenji Kawaguchi et al.ICML 2023 · 4 citations
Related papers
- SeTformer Is What You Need for Vision and LanguagePourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger, Michael FelsbergAAAI 2024 · 8 citations
- A Trainable Optimal Transport Embedding for Feature Aggregation and its Relationship to AttentionGrégoire Mialon, Dexiong Chen, Alexandre d'Aspremont, Julien MairalICLR 2021 · 71 citations
- Statistical Optimal Transport posed as Learning Kernel EmbeddingJagarlapudi Saketha Nath, Pratik Kumar JawanpuriaNeurIPS 2020 · 18 citations
- Learning Prototype-oriented Set Representations for Meta-LearningDandan Guo, Long Tian, Minghe Zhang, Mingyuan Zhou et al.ICLR 2022 · 27 citations
- A VAE for Transformers with Nonparametric Variational Information BottleneckJames Henderson, Fabio FehrICLR 2023
