Towards scientific discovery with dictionary learning: Extracting biological concepts from microscopy foundation models
Konstantin Donhauser, Kristina Ulicna, Gemma E. Moran, Aditya Ravuri, Kian Kenyon-Dean, Cian Eastwood, Jason S. Hartford
Abstract
Dictionary learning (DL) has emerged as a powerful interpretability tool for large language models. By extracting known concepts (e.g., Golden-Gate Bridge) from human-interpretable data (e.g., text), sparse DL can elucidate a model's inner workings. In this work, we ask if DL can also be used to discover unknown concepts from less human-interpretable scientific data (e.g., cell images), ultimately enabling modern approaches to scientific discovery. As a first step, we use DL algorithms to study microscopy foundation models trained on multi-cell image data, where little prior knowledge exists regarding which high-level concepts should arise. We show that sparse dictionaries indeed extract biologically-meaningful concepts such as cell type and genetic perturbation type. We also propose a new DL algorithm, Iterative Codebook Feature Learning (ICFL) and combine it with a pre-processing step which uses PCA whitening from a control dataset. In our experiments, we demonstrate that both ICFL and PCA improve the selectivity or "monosemanticity" of extracted features compared to TopK sparse autoencoders.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Toward Identifiable Sparse AutoencodersWalter Nelson, Theofanis Karaletsos, Francesco LocatelloICML 2026 · 1 citation
- The Perception–Physics Paradox: Probing Scientific Alignment with TC-BenchDingling Yao, Andrea Polesello, Adeel Pervez, Caroline Muller et al.ICML 2026
- From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language ModelsEtowah Adams, Liam Bai, Minji Lee, Yiyang Yu et al.ICML 2025
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- Codebook Features: Sparse and Discrete Interpretability for Neural NetworksAlex Tamkin, Mohammad Taufeeque, Noah D. GoodmanICML 2024 · 42 citations
- Scaling and evaluating sparse autoencodersLeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh et al.ICLR 2025 · 10 citations
- Exemplary Natural Images Explain CNN Activations Better than State-of-the-Art Feature VisualizationJudy Borowski, Roland Simon Zimmermann, Judith Schepers, Robert Geirhos et al.ICLR 2021 · 5 citations
Related papers
- Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary LearningJohn Wu, David Wu, Jimeng SunEMNLP 2024 · 4 citations
- A Concept-Based Explainability Framework for Large Multimodal ModelsJayneel Parekh, Pegah Khayatan, Mustafa Shukor, Alasdair Newson et al.NeurIPS 2024 · 48 citations
- Towards Principled Evaluations of Sparse Autoencoders for Interpretability and ControlAleksandar Makelov, Georg Lange, Neel NandaICLR 2025
- Constructing Interpretable Features from Compositional Neuron GroupsOr David Shafran, Atticus Geiger, Mor GevaACL 2026 · 4 citations
- AbsTopK: Rethinking Sparse Autoencoders For Bidirectional FeaturesXudong Zhu, Mohammad Mahdi Khalili, Zhihui ZhuICLR 2026 · 10 citations
