ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse Autoencoders
Xiangyu Liu, Haodi Lei, Yi Liu, Yang Liu, Wei Hu
Abstract
Sparse Autoencoder (SAE) has emerged as a powerful tool for mechanistic interpretability of large language models. Recent works apply SAE to protein language models (PLMs), aiming to extract and analyze biologically meaningful features from their latent spaces. However, SAE suffers from semantic entanglement, where individual neurons often mix multiple nonlinear concepts, making it difficult to reliably interpret or manipulate model behaviors. In this paper, we propose a semantically-guided SAE, called ProtSAE. Unlike existing SAE which requires annotation datasets to filter and interpret activations, we guide semantic disentanglement during training using both annotation datasets and domain knowledge to mitigate the effects of entangled attributes. We design interpretability experiments showing that ProtSAE learns more biologically relevant and interpretable hidden features compared to previous methods. Performance analyses further demonstrate that ProtSAE maintains high reconstruction fidelity while achieving better results in interpretable probing. We also show the potential of ProtSAE in steering PLMs for downstream generation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c43e363-5d55-4994-8ff2-06925733dd26Cited by top-tier papers1
Ask how each one uses itBuilds on16
- Transformer protein language models are unsupervised structure learnersRoshan Rao, Joshua Meier, Tom Sercu, Sergey Ovchinnikov et al.ICLR 2021 · 366 citations
- BERTology Meets Biology: Interpreting Attention in Protein Language ModelsJesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong et al.ICLR 2021 · 357 citations
- Function Vectors in Large Language ModelsEric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller et al.ICLR 2024 · 229 citations
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 222 citations
- Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)Usha Bhalla, Alex Oesterling, Suraj Srinivas, Flávio P. Calmon et al.NeurIPS 2024 · 146 citations
Related papers
- On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted RemedyJingyi Cui, Qi Zhang, Yifei Wang, Yisen WangICLR 2026 · 16 citations
- Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse AutoencodersDavid Chanin, Adrià Garriga-AlonsoICML 2026 · 8 citations
- Step-Level Sparse Autoencoder for Reasoning Process InterpretationXuan Yang, Jiayu Liu, Yuhang Lai, Hao Xu et al.ICML 2026 · 2 citations
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersDavid Chanin, James Wilken-Smith, Tomás Dulka, Hardik Bhatnagar et al.NeurIPS 2025 · 168 citations
- From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language ModelsEtowah Adams, Liam Bai, Minji Lee, Yiyang Yu et al.ICML 2025
