Protein Fold Classification at Scale: Benchmarking and Pretraining
Dexiong Chen, Andrei Manolache, Mathias Niepert, Karsten Borgwardt
摘要
Classifying protein topology is essential for deciphering biological function, but progress is held back by the lack of large-scale benchmarks that avoid duplicates and by models that do not scale well. We introduce TEDBench, a large-scale, nonredundant benchmark for protein fold classification constructed from the Encyclopedia of Domains (TED) and Foldseek-clustered AlphaFold structures. We show that on TEDBench, current protein representation learning methods either require very large models or fail to deliver strong performance. To address this challenge, we propose Masked Invariant Autoencoders (MiAE), a self-supervised framework for protein structure representation learning. MiAE uses an extremely high masking ratio of up to 90% with an SE(3)invariant encoder and a lightweight decoder that reconstructs backbone coordinates from the latent representation and mask tokens. MiAE scales well and outperforms supervised counterparts and state-of-the-art baselines on TEDBench, establishing a strong recipe for protein fold classification. To test transfer beyond AlphaFold structures, we further benchmark on a curated dataset from experimental structures of CATH v4.4. TED-Bench is available at https://github.com/ BorgwardtLab/TEDBench .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Protein Representation Learning by Geometric Structure PretrainingZuobai Zhang, Minghao Xu, Arian Rokkum Jamasb, Vijil Chenthamarakshan 等ICLR 2023 · 被引用 40 次
- GotenNet: Rethinking Efficient 3D Equivariant Graph Neural NetworksSarp Aykent, Tian XiaICLR 2025
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li 等CVPR 2022
相关 Paper
- Evaluating Representation Learning on the Protein Structure UniverseArian Rokkum Jamasb, Alex Morehead, Chaitanya K. Joshi, Zuobai Zhang 等ICLR 2024 · 被引用 26 次
- AlphaFold Database Debiasing for Robust Inverse FoldingCheng Tan, Zhenxiao Cao, Zhangyang Gao, Siyuan Li 等NeurIPS 2025 · 被引用 3 次
- Self-supervised learning of Split Invariant Equivariant representationsQuentin Garrido, Laurent Najman, Yann LeCunICML 2023 · 被引用 43 次
- Self-Guided Masked Autoencoders for Domain-Agnostic Self-Supervised LearningJohnathan Xie, Yoonho Lee, Annie S. Chen, Chelsea FinnICLR 2024 · 被引用 4 次
- Self-Guided Masked AutoencoderJeongwoo Shin, Inseo Lee, Junho Lee, Joonseok LeeNeurIPS 2024 · 被引用 18 次
