GlycanML: A Multi-Task and Multi-Structure Benchmark for Glycan Machine Learning
Minghao Xu, Yunteng Geng, Yihang Zhang, Ling Yang, Jian Tang, Wentao Zhang
摘要
Glycans are basic biomolecules and perform essential functions within living organisms. The rapid increase of functional glycan data provides a good opportunity for machine learning solutions to glycan understanding. However, there still lacks a standard machine learning benchmark for glycan property and function prediction. In this work, we fill this blank by building a comprehensive benchmark for Glycan Machine Learning (GLYCANML). The GLYCANML benchmark consists of diverse types of tasks including glycan taxonomy prediction, glycan immunogenicity prediction, glycosylation type prediction, and protein-glycan interaction prediction. Glycans can be represented by both sequences and graphs in GLYCANML, which enables us to extensively evaluate sequence-based models and graph neural networks (GNNs) on benchmark tasks. Furthermore, by concurrently performing eight glycan taxonomy prediction tasks, we introduce the GLYCANML-MTL testbed for multi-task learning (MTL) algorithms. Also, we evaluate how taxonomy prediction can boost other three function prediction tasks by MTL. Experimental results show the superiority of modeling glycans with multi-relational GNNs, and suitable MTL methods can further boost model performance. We provide all datasets and source codes at https://github.com/GlycanML/GlycanML and maintain a leaderboard at https://GlycanML.github.io/project .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng 等NeurIPS 2021 · 被引用 1,632 次
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu 等NeurIPS 2022 · 被引用 1,216 次
- Composition-based Multi-Relational Graph Convolutional NetworksShikhar Vashishth, Soumya Sanyal, Vikram Nitin, Partha P. TalukdarICLR 2020 · 被引用 1,105 次
- Protein Representation Learning by Geometric Structure PretrainingZuobai Zhang, Minghao Xu, Arian Rokkum Jamasb, Vijil Chenthamarakshan 等ICLR 2023 · 被引用 40 次
相关 Paper
- Modeling All-Atom Glycan Structures via Hierarchical Message Passing and Multi-Scale Pre-trainingMinghao Xu, Jiaze Song, Keming Wu, Xiangxin Zhou 等ICML 2025
- HeMeNet: Heterogeneous Multichannel Equivariant Network for Protein Multi-task LearningRong Han, Wenbing Huang, Lingxiao Luo, Xinyan Han 等AAAI 2025
- Learning Over Molecular Conformer Ensembles: Datasets and BenchmarksYanqiao Zhu, Jeehyun Hwang, Keir Adams, Zhen Liu 等ICLR 2024 · 被引用 13 次
- PepBenchmark: A Standardized Benchmark for Peptide Machine LearningJiahui Zhang, Rouyi Wang, Kuangqi Zhou, Tianshu Xiao 等ICLR 2026 · 被引用 11 次
- THGB: A Comprehensive Benchmark for Text-attributed Heterogeneous GraphsLixin Zhou, Zemin Liu, Yuan Fang, Dan Niu 等AAAI 2026
