MAESTER: Masked Autoencoder Guided Segmentation at Pixel Resolution for Accurate, Self-Supervised Subcellular Structure Recognition
Ronald Xie, Kuan Pang, Gary D. Bader, Bo Wang
Abstract
Accurate segmentation of cellular images remains an elusive task due to the intrinsic variability in morphology of biological structures. Complete manual segmentation is unfeasible for large datasets, and while supervised methods have been proposed to automate segmentation, they often rely on manually generated ground truths which are especially challenging and time consuming to generate in biology due to the requirement of domain expertise. Furthermore, these methods have limited generalization capacity, requiring additional manual labels to be generated for each dataset and use case. We introduce MAESTER (Masked AutoEncoder guided SegmenTation at pixEl Resolution), a self-supervised method for accurate, subcellular structure segmentation at pixel resolution. MAESTER treats segmentation as a representation learning and clustering problem. Specifically, MAESTER learns semantically meaningful token representations of multi-pixel image patches while simultaneously maintaining a sufficiently large field of view for contextual learning. We also develop a cover-and-stride inference strategy to achieve pixel-level subcellular structure segmentation. We evaluated MAESTER on a publicly available volumetric electron microscopy (VEM) dataset of primary mouse pancreatic islets β cells and achieved upwards of 29.1% improvement over state-of-the-art under the same evaluation criteria. Furthermore, our results are competitive against supervised methods trained on the same tasks, closing the gap between self-supervised and supervised approaches. MAESTER shows promise for alleviating the critical bottleneck of ground truth generation for imaging related data analysis and thereby greatly increasing the rate of biological discovery. Code available at https : / / github . com / bowang-lab/MAESTER
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Electron Microscopy Images as Set of Fragments for Mitochondrial SegmentationNaisong Luo, Rui Sun, Yuwen Pan, Tianzhu Zhang et al.AAAI 2024 · 9 citations
- SCCS: Deep Neural Spectral Clustering for Self-Supervised Subcellular Structure SegmentationJimao Jiang, Diya Sun, Tianbing Wang, Yuru PeiAAAI 2025 · 1 citation
- SCE-MAE: Selective Correspondence Enhancement with Masked Autoencoder for Self-Supervised Landmark EstimationKejia Yin, Varshanth S. Rao, Ruowei Jiang, Xudong Liu et al.CVPR 2024 · 1 citation
- ε-Seg: Sparsely Supervised Semantic Segmentation of Microscopy DataSheida Rahnamai Kordasiabi, Damian Dalle Nogare, Florian JugNeurIPS 2025
Builds on7
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
Related papers
- Self-Guided Masked AutoencoderJeongwoo Shin, Inseo Lee, Junho Lee, Joonseok LeeNeurIPS 2024 · 18 citations
- Masked Autoencoders for Microscopy are Scalable Learners of Cellular BiologyOren Kraus, Kian Kenyon-Dean, Saber Saberian, Maryam Fallah et al.CVPR 2024
- Unsupervised Learning of Object-Centric Embeddings for Cell Instance Segmentation in Microscopy ImagesSteffen Wolf, Manan Lalit, Katie McDole, Jan FunkeICCV 2023 · 11 citations
- R-MAE: Regions Meet Masked AutoencodersDuy-Kien Nguyen, Yanghao Li, Vaibhav Aggarwal, Martin R. Oswald et al.ICLR 2024 · 18 citations
- MAESTRO: Masked Encoding Set Transformer with Self-DistillationMatthew Eric Lee, Jaesik Kim, Matei Ionita, Jonghyun Lee et al.ICLR 2025
