ViTally Consistent: Scaling Biological Representation Learning for Cell Microscopy
Kian Kenyon-Dean, Zitong Jerry Wang, John Urbanik, Konstantin Donhauser, Jason S. Hartford, Saber Saberian, Nil Sahin, Ihab Bendidi, Safiye Celik, Juan Sebastián Rodríguez Vera, Marta M. Fay, Imran S. Haque, Oren Kraus
摘要
Deriving insights from experimentally generated datasets requires methods that can account for random and systematic measurement errors and remove them in order to accurately represent the underlying effects of the conditions being tested. Here we present a framework for pretraining on large-scale microscopy datasets that includes three steps: (1) curating a set of diverse and selfconsistent training samples, (2) scaling training of an appropriate foundation model architecture on this dataset, (3) evaluating intermediate layers of the trained model to identify the best representation for downstream tasks. Using this strategy, we present the largest foundation model for cell microscopy data to our knowledge, a new 1.9 billion-parameter ViT-G/8 MAE trained on over 8 billion microscopy image crops. Compared to a previous published ViT-L/8 MAE, our new model achieves a 60% improvement in linear separability of genetic perturbations and obtains the best overall performance on whole-genome relationship recall, batch correction replicate consistency, and compound-gene activity prediction benchmarks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- ChA-MAEViT: Unifying Channel-Aware Masked Autoencoders and Multi-Channel Vision Transformers for Improved Cross-Channel LearningChau Pham, Juan C. Caicedo, Bryan A. PlummerNeurIPS 2025 · 被引用 11 次
- CHAMMI-75: Pre-training multi-channel models with heterogeneous microscopy imagesVidit Agrawal, John Peters, Tyler N. Thompson, Mohammad V. Sanian 等ICLR 2026 · 被引用 3 次
- Asymmetric Contrastive Objectives for Efficient Phenotypic ScreeningLuke Nightingale, Joseph Tuersley, Scott Warchal, Andrea Cairoli 等ICML 2026 · 被引用 1 次
- Towards scientific discovery with dictionary learning: Extracting biological concepts from microscopy foundation modelsKonstantin Donhauser, Kristina Ulicna, Gemma E. Moran, Aditya Ravuri 等ICML 2025
- Deep Learning for BioImaging: What Are We Really Learning?Ivan Svatko, Maxime Sanchez, Ihab Bendidi, Gilles Cottrell 等ICML 2026
它引用的顶会 Paper12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real 等NeurIPS 2023 · 被引用 734 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
相关 Paper
- Masked Autoencoders for Microscopy are Scalable Learners of Cellular BiologyOren Kraus, Kian Kenyon-Dean, Saber Saberian, Maryam Fallah 等CVPR 2024
- Integrating Biological Knowledge for Robust Microscopy Image Profiling on De Novo Cell LinesJiayuan Chen, Thai-Hoang Pham, Yuanlong Wang, Ping ZhangICCV 2025
- MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in MicroscopyAlbert Dominguez Mantes, Gioele La Manno, Martin WeigertCVPR 2026 · 被引用 1 次
- EVA: Exploring the Limits of Masked Visual Representation Learning at ScaleYuxin Fang, Wen Wang, Binhui Xie, Quan Sun 等CVPR 2023
- VideoMAE V2: Scaling Video Masked Autoencoders with Dual MaskingLimin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong 等CVPR 2023
