Yet Another ICU Benchmark: A Flexible Multi-Center Framework for Clinical ML
Robin Van De Water, Hendrik Schmidt, Paul W. G. Elbers, Patrick Thoral, Bert Arnrich, Patrick Rockenschaub
Abstract
Medical applications of machine learning (ML) have experienced a surge in popularity in recent years. The intensive care unit (ICU) is a natural habitat for ML given the abundance of available data from electronic health records. Models have been proposed to address numerous ICU prediction tasks like the early detection of complications. While authors frequently report state-of-the-art performance, it is challenging to verify claims of superiority. Datasets and code are not always published, and cohort definitions, preprocessing pipelines, and training setups are difficult to reproduce. This work introduces Yet Another ICU Benchmark (YAIB), a modular framework that allows researchers to define reproducible and comparable clinical ML experiments; we offer an end-to-end solution from cohort definition to model evaluation. The framework natively supports most open-access ICU datasets (MIMIC III/IV, eICU, HiRID, AUMCdb) and is easily adaptable to future ICU datasets. Combined with a transparent preprocessing pipeline and extensible training code for multiple ML and deep learning models, YAIB enables unified model development. Our benchmark comes with five predefined established prediction tasks (mortality, acute kidney injury, sepsis, kidney function, and length of stay) developed in collaboration with clinicians. Adding further tasks is straightforward by design. Using YAIB, we demonstrate that the choice of dataset, cohort definition, and preprocessing have a major impact on the prediction performance - often more so than model class - indicating an urgent need for YAIB as a holistic benchmarking tool. We provide our work to the clinical ML community to accelerate method development and enable real-world clinical implementations. Software Repository: https://github.com/rvandewater/YAIB.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- FEDKIM: Adaptive Federated Knowledge Injection into Medical Foundation ModelsXiaochen Wang, Jiaqi Wang, Houping Xiao, Jinghui Chen et al.EMNLP 2024 · 6 citations
- Can we generate portable representations for clinical time series data using LLMs?Zongliang Ji, Yifei Sun, Andre Carlos Kajdacsy-Balla Amaral, Anna Goldenberg et al.ICLR 2026 · 5 citations
- PyHealth 2.0: A Comprehensive Open-Source Toolkit for Accessible and Reproducible Clinical Deep LearningJohn Wu, Yongda Fan, Zhenbang Wu, Paul Landes et al.ICML 2026 · 1 citation
- ACES: Automatic Cohort Extraction System for Event-Stream DatasetsJustin Xu, Jack Gallifant, Alistair E. W. Johnson, Matthew B. A. McDermottICLR 2025
- Benchmarking Reinforcement Learning Algorithms for ICU Ventilator Settings: An Interpretable and Probabilistic Patient Environment for Doctor AgentsYa-Hsi Chang, Po-Chih KuoAAAI 2026
Builds on4
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- CSDI: Conditional Score-based Diffusion Models for Probabilistic Time Series ImputationYusuke Tashiro, Jiaming Song, Yang Song, Stefano ErmonNeurIPS 2021 · 1,245 citations
- HyperImpute: Generalized Iterative Imputation with Automatic Model SelectionDaniel Jarrett, Bogdan Cebere, Tennison Liu, Alicia Curth et al.ICML 2022 · 129 citations
Related papers
- Distilling Knowledge from Publicly Available Online EMR Data to Emerging Epidemic for PrognosisLiantao Ma, Xinyu Ma, Junyi Gao, Xianfeng Jiao et al.WWW 2021 · 32 citations
- CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation ModelsWei Dai, Peilin Chen, Malinda Lu, Daniel Li et al.ICML 2025
- REACT-LLM: A Benchmark for Evaluating LLM Integration with Causal Features in Clinical Prognostic TasksLinna Wang, Zhixuan You, Qihui Zhang, Jiunan Wen et al.AAAI 2026
- Clairvoyance: A Pipeline Toolkit for Medical Time SeriesDaniel Jarrett, Jinsung Yoon, Ioana Bica, Zhaozhi Qian et al.ICLR 2021 · 43 citations
- Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community RetrievalPengcheng Jiang, Cao Xiao, Minhao Jiang, Parminder Bhatia et al.ICLR 2025
