Protein Circuit Tracing via Cross-layer Transcoders
Darin Tsui, Kunal Talreja, Daniel Saeedi, Amirali Aghazadeh
Abstract
Protein language models (pLMs) have emerged as powerful predictors of protein structure and function. However, the computational circuits underlying their predictions remain poorly understood. Recent mechanistic interpretability methods decompose pLM representations into interpretable features, but they treat each layer independently and thus fail to capture cross-layer computation, limiting their ability to approximate the full model. We introduce ProtoMech, a framework for discovering computational circuits in pLMs using cross-layer transcoders that learn sparse latent representations jointly across layers to capture the model’s full computational circuitry. Applied to the pLM ESM2, ProtoMech recovers 82–89% of the original performance on protein family classification and function prediction tasks. ProtoMech then identifies compressed circuits that use <1% of the latent space while retaining up to 79% of model accuracy, revealing correspondence with structural and functional motifs, including binding, signaling, and stability. Steering along these circuits enables high-fitness protein design, surpassing baseline methods in more than 70% of cases. These results establish ProtoMech as a principled framework for protein circuit tracing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 222 citations
- Scaling Unlocks Broader Generation and Deeper Functional Understanding of ProteinsAadyot Bhatnagar, Sarthak Jain, Joel Beazer, Samuel Curran et al.NeurIPS 2025 · 63 citations
- Proximal Exploration for Model-guided Protein Sequence DesignZhizhou Ren, Jiahan Li, Fan Ding, Yuan Zhou et al.ICML 2022 · 52 citations
- ProteinNPT: Improving protein property prediction and design with non-parametric transformersPascal Notin, Ruben Weitzman, Debora S. Marks, Yarin GalNeurIPS 2023 · 48 citations
- Improving protein optimization with smoothed fitness landscapesAndrew Kirjner, Jason Yim, Raman Samusevich, Shahar Bracha et al.ICLR 2024 · 28 citations
Related papers
- From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language ModelsEtowah Adams, Liam Bai, Minji Lee, Yiyang Yu et al.ICML 2025
- ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse AutoencodersXiangyu Liu, Haodi Lei, Yi Liu, Yang Liu et al.AAAI 2026 · 2 citations
- Beyond Components: Singular Vector-Based Interpretability of Transformer CircuitsAreeb Ahmad, Abhinav Joshi, Ashutosh ModiNeurIPS 2025 · 9 citations
- Towards Understanding the Shape of Representations in Protein Language ModelsKosio Beshkov, Anders Malthe-SørenssenICLR 2026 · 2 citations
- Weight-sparse transformers have interpretable circuitsLeo Gao, Achyuta Rajaram, Jacob Coxon, Soham Govande et al.ICML 2026
