Protein Fold Classification at Scale: Benchmarking and Pretraining
Dexiong Chen, Andrei Manolache, Mathias Niepert, Karsten Borgwardt
Abstract
Classifying protein topology is essential for deciphering biological function, but progress is held back by the lack of large-scale benchmarks that avoid duplicates and by models that do not scale well. We introduce TEDBench, a large-scale, nonredundant benchmark for protein fold classification constructed from the Encyclopedia of Domains (TED) and Foldseek-clustered AlphaFold structures. We show that on TEDBench, current protein representation learning methods either require very large models or fail to deliver strong performance. To address this challenge, we propose Masked Invariant Autoencoders (MiAE), a self-supervised framework for protein structure representation learning. MiAE uses an extremely high masking ratio of up to 90% with an SE(3)invariant encoder and a lightweight decoder that reconstructs backbone coordinates from the latent representation and mask tokens. MiAE scales well and outperforms supervised counterparts and state-of-the-art baselines on TEDBench, establishing a strong recipe for protein fold classification. To test transfer beyond AlphaFold structures, we further benchmark on a curated dataset from experimental structures of CATH v4.4. TED-Bench is available at https://github.com/ BorgwardtLab/TEDBench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf1eee80-24ae-4c73-929b-901cb3d65c91Builds on4
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Protein Representation Learning by Geometric Structure PretrainingZuobai Zhang, Minghao Xu, Arian Rokkum Jamasb, Vijil Chenthamarakshan et al.ICLR 2023 · 40 citations
- GotenNet: Rethinking Efficient 3D Equivariant Graph Neural NetworksSarp Aykent, Tian XiaICLR 2025
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li et al.CVPR 2022
Related papers
- Evaluating Representation Learning on the Protein Structure UniverseArian Rokkum Jamasb, Alex Morehead, Chaitanya K. Joshi, Zuobai Zhang et al.ICLR 2024 · 26 citations
- AlphaFold Database Debiasing for Robust Inverse FoldingCheng Tan, Zhenxiao Cao, Zhangyang Gao, Siyuan Li et al.NeurIPS 2025 · 3 citations
- Self-supervised learning of Split Invariant Equivariant representationsQuentin Garrido, Laurent Najman, Yann LeCunICML 2023 · 43 citations
- Self-Guided Masked Autoencoders for Domain-Agnostic Self-Supervised LearningJohnathan Xie, Yoonho Lee, Annie S. Chen, Chelsea FinnICLR 2024 · 4 citations
- Self-Guided Masked AutoencoderJeongwoo Shin, Inseo Lee, Junho Lee, Joonseok LeeNeurIPS 2024 · 18 citations
