Quantifying the Generalization Gap in Seizure Detection: A Large-Scale Empirical Benchmark via the SzCORE Challenge
Jonathan Dan, Amirhossein Shahbazinia, Christodoulos Kechris, David Atienza
摘要
Reliable automatic seizure detection from long-term electroencephalogram recordings (EEG) remains an unsolved challenge, as current models often fail to generalize across patients or clinical settings. Manual EEG review still is the standard of care, highlighting the need for robust models and standardized evaluation. The current literature often reports high efficacy, yet these models frequently fail when deployed to unseen patient populations. To rigorously assess this generalization gap, we conducted a large-scale empirical study evaluating 28 state-of-the-art algorithmic architectures, ranging from classical feature engineering to modern Deep Learning. These algorithms were collected by organizing a competition. A strictly held-out private dataset of continuous EEG recordings from 65 subjects, totaling 4,360 hours of data, was utilized to evaluate algorithm performance. Expert neurophysiologists annotated these recordings, establishing the ground truth for seizure events. Algorithms were evaluated using event-based metrics from the SzCORE framework, including sensitivity, precision, F1-score, and false positive rate per day. Results revealed significant performance variability among state-of-the-art approaches, with the top F1 score of 32% (sensitivity 37%, precision 29%), highlighting the persistent difficulty of this task for current machine learning methodologies. Our analysis uncovered a discordance between peak performance and population-level stability. The algorithms achieving the highest aggregate F1-scores did not achieve the most consistent ranking across subjects, indicating high performance variance and susceptibility to failure on outlier patients. This independent evaluation also exposed a notable gap between self-reported efficacies and hold-out performance, underscoring the critical need for standardized, rigorous benchmarking in developing clinically viable ML models. A comparison with previous challenges and commercial systems indicates that the best algorithm in this study surpassed prior methods. Critically, the evaluation infrastructure transitions into a continuously open benchmarking platform, fostering reproducible research and accelerating the development of robust seizure detection algorithms by allowing ongoing submissions and integration of additional private datasets. Clinical centers can also adopt this platform to evaluate seizure detection algorithms on their EEG data using a standardized, reproducible framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCIWei-Bang Jiang, Li-Ming Zhao, Bao-Liang LuICLR 2024 · 被引用 298 次
- Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitectureMahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski 等CVPR 2023
相关 Paper
- DMNet: Self-comparison Driven Model for Subject-independent Seizure DetectionShihao Tu, Linfeng Cao, Daoze Zhang, Junru Chen 等NeurIPS 2024 · 被引用 6 次
- Omni-iEEG: A Large-Scale, Comprehensive iEEG Dataset and Benchmark for Epilepsy ResearchChenda Duan, Yipeng Zhang, Sotaro Kanai, Yuanyi Ding 等ICLR 2026 · 被引用 7 次
- Self-Supervised Learning for Anomalous Channel Detection in EEG Graphs: Application to Seizure AnalysisThi Kieu Khanh Ho, Narges ArmanfardAAAI 2023 · 被引用 58 次
- Self-Supervised Graph Neural Networks for Improved Electroencephalographic Seizure AnalysisSiyi Tang, Jared Dunnmon, Khaled Kamal Saab, Xuan Zhang 等ICLR 2022 · 被引用 157 次
- iEDeaL: A Deep Learning Framework for Detecting Highly Imbalanced Interictal Epileptiform DischargesQitong Wang, Stephen Whitmarsh, Vincent Navarro, Themis PalpanasVLDB 2023 · 被引用 15 次
