Quantifying the Generalization Gap in Seizure Detection: A Large-Scale Empirical Benchmark via the SzCORE Challenge
Jonathan Dan, Amirhossein Shahbazinia, Christodoulos Kechris, David Atienza
Abstract
Reliable automatic seizure detection from long-term electroencephalogram recordings (EEG) remains an unsolved challenge, as current models often fail to generalize across patients or clinical settings. Manual EEG review still is the standard of care, highlighting the need for robust models and standardized evaluation. The current literature often reports high efficacy, yet these models frequently fail when deployed to unseen patient populations. To rigorously assess this generalization gap, we conducted a large-scale empirical study evaluating 28 state-of-the-art algorithmic architectures, ranging from classical feature engineering to modern Deep Learning. These algorithms were collected by organizing a competition. A strictly held-out private dataset of continuous EEG recordings from 65 subjects, totaling 4,360 hours of data, was utilized to evaluate algorithm performance. Expert neurophysiologists annotated these recordings, establishing the ground truth for seizure events. Algorithms were evaluated using event-based metrics from the SzCORE framework, including sensitivity, precision, F1-score, and false positive rate per day. Results revealed significant performance variability among state-of-the-art approaches, with the top F1 score of 32% (sensitivity 37%, precision 29%), highlighting the persistent difficulty of this task for current machine learning methodologies. Our analysis uncovered a discordance between peak performance and population-level stability. The algorithms achieving the highest aggregate F1-scores did not achieve the most consistent ranking across subjects, indicating high performance variance and susceptibility to failure on outlier patients. This independent evaluation also exposed a notable gap between self-reported efficacies and hold-out performance, underscoring the critical need for standardized, rigorous benchmarking in developing clinically viable ML models. A comparison with previous challenges and commercial systems indicates that the best algorithm in this study surpassed prior methods. Critically, the evaluation infrastructure transitions into a continuously open benchmarking platform, fostering reproducible research and accelerating the development of robust seizure detection algorithms by allowing ongoing submissions and integration of additional private datasets. Clinical centers can also adopt this platform to evaluate seizure detection algorithms on their EEG data using a standardized, reproducible framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3da616d6-4f58-4985-9917-4ca2cb6ff852Builds on3
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCIWei-Bang Jiang, Li-Ming Zhao, Bao-Liang LuICLR 2024 · 298 citations
- Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitectureMahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski et al.CVPR 2023
Related papers
- DMNet: Self-comparison Driven Model for Subject-independent Seizure DetectionShihao Tu, Linfeng Cao, Daoze Zhang, Junru Chen et al.NeurIPS 2024 · 6 citations
- Omni-iEEG: A Large-Scale, Comprehensive iEEG Dataset and Benchmark for Epilepsy ResearchChenda Duan, Yipeng Zhang, Sotaro Kanai, Yuanyi Ding et al.ICLR 2026 · 7 citations
- Self-Supervised Learning for Anomalous Channel Detection in EEG Graphs: Application to Seizure AnalysisThi Kieu Khanh Ho, Narges ArmanfardAAAI 2023 · 58 citations
- Self-Supervised Graph Neural Networks for Improved Electroencephalographic Seizure AnalysisSiyi Tang, Jared Dunnmon, Khaled Kamal Saab, Xuan Zhang et al.ICLR 2022 · 157 citations
- iEDeaL: A Deep Learning Framework for Detecting Highly Imbalanced Interictal Epileptiform DischargesQitong Wang, Stephen Whitmarsh, Vincent Navarro, Themis PalpanasVLDB 2023 · 15 citations
