Protein Circuit Tracing via Cross-layer Transcoders
Darin Tsui, Kunal Talreja, Daniel Saeedi, Amirali Aghazadeh
摘要
Protein language models (pLMs) have emerged as powerful predictors of protein structure and function. However, the computational circuits underlying their predictions remain poorly understood. Recent mechanistic interpretability methods decompose pLM representations into interpretable features, but they treat each layer independently and thus fail to capture cross-layer computation, limiting their ability to approximate the full model. We introduce ProtoMech, a framework for discovering computational circuits in pLMs using cross-layer transcoders that learn sparse latent representations jointly across layers to capture the model’s full computational circuitry. Applied to the pLM ESM2, ProtoMech recovers 82–89% of the original performance on protein family classification and function prediction tasks. ProtoMech then identifies compressed circuits that use <1% of the latent space while retaining up to 79% of model accuracy, revealing correspondence with structural and functional motifs, including binding, signaling, and stability. Steering along these circuits enables high-fitness protein design, surpassing baseline methods in more than 70% of cases. These results establish ProtoMech as a principled framework for protein circuit tracing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 被引用 222 次
- Scaling Unlocks Broader Generation and Deeper Functional Understanding of ProteinsAadyot Bhatnagar, Sarthak Jain, Joel Beazer, Samuel Curran 等NeurIPS 2025 · 被引用 63 次
- Proximal Exploration for Model-guided Protein Sequence DesignZhizhou Ren, Jiahan Li, Fan Ding, Yuan Zhou 等ICML 2022 · 被引用 52 次
- ProteinNPT: Improving protein property prediction and design with non-parametric transformersPascal Notin, Ruben Weitzman, Debora S. Marks, Yarin GalNeurIPS 2023 · 被引用 48 次
- Improving protein optimization with smoothed fitness landscapesAndrew Kirjner, Jason Yim, Raman Samusevich, Shahar Bracha 等ICLR 2024 · 被引用 28 次
相关 Paper
- From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language ModelsEtowah Adams, Liam Bai, Minji Lee, Yiyang Yu 等ICML 2025
- ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse AutoencodersXiangyu Liu, Haodi Lei, Yi Liu, Yang Liu 等AAAI 2026 · 被引用 2 次
- Beyond Components: Singular Vector-Based Interpretability of Transformer CircuitsAreeb Ahmad, Abhinav Joshi, Ashutosh ModiNeurIPS 2025 · 被引用 9 次
- Towards Understanding the Shape of Representations in Protein Language ModelsKosio Beshkov, Anders Malthe-SørenssenICLR 2026 · 被引用 2 次
- Weight-sparse transformers have interpretable circuitsLeo Gao, Achyuta Rajaram, Jacob Coxon, Soham Govande 等ICML 2026
