Cyto-SSL: A Self-Supervised Pretraining Framework for Cytology Foundation Model
Yiming Zhang, Rui Yan, Xiaohua Wan, Yifan Zhao, Shuang Feng, Zhetao Xu, Ying Wang, Fa Zhang, Bin Hu
Abstract
Cytological images originate from exfoliated cells, collected via liquid-based slides and digitized into whole slide images (WSIs). Unlike histological WSIs that exhibit continuous and well-structured tissue, cytological WSIs are sparse in spatial distribution and unstructured in cellular relationships. Typically, the nucleus serves as the primary diagnostic feature, while surrounding cytoplasmic information plays a supportive role. These unique characteristics limit the development of effective foundation models and hinder the transferability of histology-based models for cytopathology. To address this, we propose Cyto-SSL, the first self-supervised pretraining framework for cytological images. It introduces Nuclei-Centered Perturbation, which highlights individual nuclei by perturbing non-nuclear regions. We also design an SR-Transformer module, which complements this by using sparse attention to concentrate on diagnostically relevant scattered cells, while iRPE helps model to capture local spatial relationships and avoids unnecessary attention to irrelevant global structures. Experimental results show that Cyto-SSL enhances performance across diverse cytological datasets and Multiple Instance Learning (MIL) methods. On a WSI-level dataset, it achieved 95.67% accuracy and outperformed ImageNet-pretrained ResNet-50 by 11.33%, demonstrating superior feature representation for cytological analysis. Additionally, Cyto-SSL modules are plug-and-play, easily integrated into other pretraining frameworks, yielding a 2.6% accuracy gain across different SSL methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image ClassificationZhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang et al.NeurIPS 2021 · 1,163 citations
- Rethinking and Improving Relative Position Encoding for Vision TransformerKan Wu, Houwen Peng, Minghao Chen, Jianlong Fu et al.ICCV 2021 · 427 citations
- A Variational Approach for Learning from Positive and Unlabeled DataHui Chen, Fangqing Liu, Yin Wang, Liyue Zhao et al.NeurIPS 2020 · 76 citations
Related papers
- Benchmarking Self-Supervised Learning on Diverse Pathology DatasetsMingu Kang, Heon Song, Seonwook Park, Donggeun Yoo et al.CVPR 2023
- MUSE: Multi-Scale Dense Self-Distillation for Nucleus Detection and ClassificationZijiang Yang, Hanqing Chao, Bokai Zhao, Yelin Yang et al.AAAI 2026 · 2 citations
- Improving Representation Learning for Histopathologic Images with Cluster ConstraintsWeiyi Wu, Chongyang Gao, Joseph DiPalma, Soroush Vosoughi et al.ICCV 2023 · 13 citations
- Cello: A Universal Cell-wise Feature Aggregation framework for Reliable Pathology Images AnalysisHengrui Lou, Weihan Li, Jiazhen Yang, Lingxiang Jia et al.ICML 2026
- Ditch the Denoiser: Emergence of Noise Robustness in Self-Supervised Learning from Data CurriculumWenquan Lu, Jiaqi Zhang, Hugues Van Assel, Randall BalestrieroNeurIPS 2025 · 5 citations
