Controllable Latent Space Augmentation for Digital Pathology
Sofiène Boutaj, Marin Scalbert, Pierre Marza, Florent Couzinie-Devy, Maria Vakalopoulou, Stergios Christodoulidis
Abstract
Whole slide image (WSI) analysis in digital pathology presents unique challenges due to the gigapixel resolution of WSIs and the scarcity of dense supervision signals. While Multiple Instance Learning (MIL) is a natural fit for slide-level tasks, training robust models requires large and diverse datasets. Even though image augmentation techniques could be utilized to increase data variability and reduce overfitting, implementing them effectively is not a trivial task. Traditional patch-level augmentation is prohibitively expensive due to the large number of patches extracted from each WSI, and existing feature-level augmentation methods lack control over transformation semantics. We introduce HistAug, a fast and efficient generative model for controllable augmentations in the latent space for digital pathology. By conditioning on explicit patch-level transformations (e.g., hue, erosion), HistAug generates realistic augmented embeddings while preserving initial semantic information. Our method allows the processing of a large number of patches in a single forward pass efficiently, while at the same time consistently improving MIL model performance. Experiments across multiple slide-level tasks and diverse organs show that HistAug outperforms existing methods, particularly in low-data regimes. Ablation studies confirm the benefits of learned transformations over noise-based perturbations and highlight the importance of uniform WSI-wise augmentation. Code is available at https://github.com/MICS-Lab/HistAug.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2fb885bf-7601-4fc5-97b5-f236fcc4d71fBuilds on4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- RankMix: Data Augmentation for Weakly Supervised Learning of Classifying Whole Slide Images with Diverse Sizes and Imbalanced CategoriesYuan-Chih Chen, Chun-Shien LuCVPR 2023
- Dual-Stream Multiple Instance Learning Network for Whole Slide Image Classification With Self-Supervised Contrastive LearningBin Li, Yin Li, Kevin W. EliceiriCVPR 2021
Related papers
- Promptable Representation Distribution Learning and Data Augmentation for Gigapixel Histopathology WSI AnalysisKunming Tang, Zhiguo Jiang, Jun Shi, Wei Wang et al.AAAI 2025 · 1 citation
- Navigating the MIL Trade-Off: Flexible Pooling for Whole Slide Image ClassificationHossein Jafarinia, Danial Hamdi, Amirhossein Alamdar, Elahe Zahiri et al.NeurIPS 2025 · 1 citation
- Flow-MIL: Constructing Highly-expressive Latent Feature Space for Whole Slide Image Classification using Normalizing FlowYingfan Ma, Bohan An, Ao Shen, Mingzhi Yuan et al.ICCV 2025 · 1 citation
- Mixture of Mini Experts: Overcoming the Linear Layer Bottleneck in Multiple Instance LearningDaniel Shao, Joel Runevic, Richard J. Chen, Drew F. K. Williamson et al.ICLR 2026 · 3 citations
- DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image ClassificationHongrun Zhang, Yanda Meng, Yitian Zhao, Yihong Qiao et al.CVPR 2022 · 402 citations
