MAESTRO: Masked Encoding Set Transformer with Self-Distillation
Matthew Eric Lee, Jaesik Kim, Matei Ionita, Jonghyun Lee, Michelle L. McKeague, Yonghyun Nam, Irene Khavin, Yidi Huang, Victoria Fang, Sokratis Apostolidis, Divij Mathew, Shwetank
Abstract
The immune system is a complex network of cells, orchestrating coordinated responses throughout the human body over a person's lifetime. Cytometry enables profiling of this network, but current approaches focus on enumerating and phenotyping immune cells, failing to quantify the system as a whole. We present MAESTRO, a self-supervised set representation learning model that generates vector representations of set-structured data, which we apply to learn immune profile representations. Unlike previous studies that learn cell-level representations, MAESTRO uses all of a sample's cells to learn a set representation. MAE-STRO leverages attention mechanisms to handle sets of variable number of cells and ensure permutation invariance, coupled with a self-distillation framework. It is capable of reconstructing immune profiles (cells) even when 90% are hidden as its training objective. We benchmark our model against existing cytometry approaches and other existing machine learning methods that have never been applied in cytometry. Our model outperforms existing approaches in retrieving celltype distributions and capturing clinically relevant features for downstream tasks such as disease diagnosis, age, sex. Code available: https://github.com/matthew-lee1/MAESTRO
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Differentiable Expectation-Maximization for Set Representation LearningMinyoung KimICLR 2022 · 18 citations
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li et al.CVPR 2022
Related papers
- MAESTER: Masked Autoencoder Guided Segmentation at Pixel Resolution for Accurate, Self-Supervised Subcellular Structure RecognitionRonald Xie, Kuan Pang, Gary D. Bader, Bo WangCVPR 2023
- MUSE: Multi-Scale Dense Self-Distillation for Nucleus Detection and ClassificationZijiang Yang, Hanqing Chao, Bokai Zhao, Yelin Yang et al.AAAI 2026 · 2 citations
- A Label Disambiguation-Based Multimodal Massive Multiple Instance Learning Approach for Immune Repertoire ClassificationFan Xu, Yu Zhao, Bingzhe Wu, Yueshan Huang et al.AAAI 2024 · 2 citations
- A Closer Look at Self-Supervised Lightweight Vision TransformersShaoru Wang, Jin Gao, Zeming Li, Xiaoqin Zhang et al.ICML 2023 · 61 citations
- MedGMAE: Gaussian Masked Autoencoders for Medical Volumetric Representation LearningXueming Fu, Fenghe Tang, Rongsheng Wang, Yingtai Li et al.ICLR 2026
