GENEB: Why Genomic Models Are Hard to Compare
Daria Ledneva, Mikhail Nuridinov, Denis Kuznetsov
摘要
Progress in genomic foundation models is difficult to assess due to fragmented benchmarks, incompatible evaluation protocols, and task-specific reporting. As a result, claims of superiority or generality across models are often not directly comparable. We introduce GENEB, a large-scale diagnostic benchmark that evaluates frozen representations from 40 genomic foundation models across 100 tasks spanning 13 functional categories under a unified probing-based protocol, including few-shot regimes. GENEB enables controlled comparison across model scale, architecture, tokenization, and pretraining data while explicitly exposing task-level trade-offs. Our analysis shows that aggregate leaderboards are unstable: model rankings vary sharply across task categories, scale provides only modest and inconsistent gains, and architectural and pretraining alignment frequently outweigh parameter count. These results highlight limitations of current evaluation practices and position GENEB as a reference framework for principled comparison and category-aware model selection in genomic machine learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence ModelingYair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao 等ICML 2024 · 被引用 195 次
- Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNALifeng Qiao, Peng Ye, Yuchen Ren, Weiqiang Bai 等NeurIPS 2024 · 被引用 23 次
- BioToken and BioFM – Biologically-Informed Tokenization Enables Accurate and Efficient Genomic Foundation ModelsAleksandr Medvedev, Karthik Viswanathan, Praveenkumar Kanithi, Kirill Vishniakov 等ICML 2026
- SPACE: Your Genomic Profile Predictor is a Powerful DNA Foundation ModelZhao Yang, Jiwei Zhu, Bing SuICML 2025
相关 Paper
- MORE: Molecule Pretraining with Multi-Level Pretext TaskYeongyeong Son, Dasom Noh, Gyoungyoung Heo, Gyoung Jin Park 等AAAI 2025 · 被引用 1 次
- Tokenization to Transfer: Do Genomic Foundation Models Learn Good Representations?Kirill Vishniakov, Karthik Viswanathan, Aleksandr Medvedev, Praveenkumar Kanithi 等ICLR 2026 · 被引用 16 次
- EEG-FM-Bench: A Comprehensive Benchmark for the Systematic Evaluation and Diagnostic Analyses of EEG Foundation ModelsWei Xiong, Jiangtong Li, Jie Li, Kun Zhu 等ICML 2026 · 被引用 15 次
- Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation ModelsWeimin Wu, Xuefeng Song, Yibo Wen, Qinjie Lin 等ICML 2026
- GenomeQA: Benchmarking General Large Language Models for Genome Sequence UnderstandingWeicai Long, Yusen Hou, Junning Feng, Houcheng Su 等ACL 2026
