VenusX: Unlocking Fine-Grained Functional Understanding of Proteins
Yang Tan, Wenrui Gou, Bozitao Zhong, Huiqun Yu, Liang Hong, Bingxin Zhou
摘要
Deep learning models have driven significant progress in predicting protein function and interactions at the protein level. While these advancements have been invaluable for many biological applications such as enzyme engineering and function annotation, a more detailed perspective is essential for understanding protein functional mechanisms and evaluating the biological knowledge captured by models. This study introduces VenusX, the first benchmark designed to assess protein representation learning with a focus on fine-grained intra-protein functional understanding. VenusX comprises three major task categories across six types of annotations, including residue-level binary classification, fragment-level multi-class classification, and pairwise functional similarity scoring for identifying critical active sites, binding sites, conserved sites, motifs, domains, and epitopes. The benchmark features over 878,000 samples curated from major open-source databases such as InterPro, BioLiP, and SAbDab. By providing mixed-family and cross-family splits at three sequence identity thresholds, our benchmark enables a comprehensive assessment of model performance on both in-distribution and out-of-distribution scenarios. For baseline evaluation, we assess a diverse set of popular and open-source models, including pre-trained protein language models, sequence-structure hybrids, structure-based methods, and alignment-based techniques. Their performance is reported across all benchmark datasets and evaluation settings using multiple metrics, offering a thorough comparison and a strong foundation for future research. Our code (https://github.com/ai4protein/VenusX), data (https://huggingface.co/collections/AI4Protein/venusx-dataset), and a leaderboard (https://ai4protein.github.io/venusx/) are provided as open-source resources.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Learning from Protein Structure with Geometric Vector PerceptronsBowen Jing, Stephan Eismann, Patricia Suriana, Raphael John Lamarre Townshend 等ICLR 2021 · 被引用 627 次
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin 等ICML 2022 · 被引用 560 次
- SaProt: Protein Language Modeling with Structure-aware VocabularyJin Su, Chenchen Han, Yuyang Zhou, Junjie Shan 等ICLR 2024 · 被引用 285 次
- Immunogenicity Prediction with Dual Attention Enables Vaccine Target SelectionSong Li, Yang Tan, Song Ke, Liang Hong 等ICLR 2025
相关 Paper
- Greater than the Sum of Its Parts: Building Substructure into Protein Encoding ModelsRobert Calef, Arthur Liang, Manolis Kellis, Marinka ZitnikICLR 2026 · 被引用 2 次
- ProtCLIP: Function-Informed Protein Multi-Modal LearningHanjing Zhou, Mingze Yin, Wei Wu, Mingyang Li 等AAAI 2025 · 被引用 11 次
- AtomSurf: Surface Representation for Learning on Protein StructuresVincent Mallet, Yangyang Miao, Souhaib Attaiki, Bruno Correia 等ICLR 2025
- PoseX: AI Defeats Physics-based Methods on Protein Ligand Cross-DockingYize Jiang, Xinze Li, Yuanyuan Zhang, Jin Han 等ICLR 2026 · 被引用 5 次
- Fast End-to-End Learning on Protein SurfacesFreyr Sverrisson, Jean Feydy, Bruno E. Correia, Michael M. BronsteinCVPR 2021
