BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics
Dionizije Fa, Marko Culjak, Bruno Pandza, Mateo Cupic
摘要
We introduce BioAgent Bench, an evaluation suite designed for measuring the performance and robustness of AI agents in common bioinformatics tasks. The suite consists of manually curated end-to-end tasks (e.g., RNA-seq, variant calling, metagenomics) accompanied by taskspecific prompts and concrete output artifacts to support automated assessment. We evaluate frontier closed-and open-weight models across multiple agent harnesses, and use an LLM-based grader to score pipeline progress and outcome validity. We find that agents based on frontier LLMs can complete multi-step bioinformatics pipelines without elaborate custom scaffolding, often producing the requested final artifacts reliably. However, robustness tests reveal failure modes under controlled perturbations (corrupted inputs, decoy files, and prompt bloat), indicating that correct high-level pipeline construction does not guarantee reliable step-level reasoning. By releasing the code and the complementary resources constituting our suite, our primary goal is to accelerate the development of cost-effective yet reliable local agents, capable of handling complex bioinformatics workflows often involving sensitive patient data or unpublished intellectual property. github/bioagent-bench github/bioagent-experiments
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu 等ICLR 2024 · 被引用 748 次
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMsMinghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song 等EMNLP 2023 · 被引用 72 次
- HeurekaBench: A Benchmarking Framework for AI Co-scientistSiba Smarak Panigrahi, Jovana Videnovic, Maria BrbicICLR 2026 · 被引用 10 次
相关 Paper
- BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous ScienceYuyang Liu, Liuzhenghao Lyu, Xiancheng Zhang, Jingya Wang 等ICML 2026
- ABC-Bench: An Agentic Bio-Capabilities Benchmark for BiosecurityAndrew Liu, Samira Nedungadi, Bryce Cai, Alex Kleinman 等ICML 2026 · 被引用 6 次
- ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT PipelinesTengjun Jin, Yuxuan Zhu, Daniel KangVLDB 2026 · 被引用 13 次
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line InterfacesMike A. Merrill, Alexander Glenn Shaw, Nicholas Carlini, Boxuan Li 等ICLR 2026 · 被引用 520 次
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang 等ACL 2026 · 被引用 14 次
