EEG-FM-Bench: A Comprehensive Benchmark for the Systematic Evaluation and Diagnostic Analyses of EEG Foundation Models
Wei Xiong, Jiangtong Li, Jie Li, Kun Zhu, Changjun Jiang
Abstract
Electroencephalography foundation models (EEG-FMs) have advanced brain signal analysis, but the lack of standardized evaluation benchmarks impedes model comparison and scientific progress. Current evaluations rely on inconsistent protocols that render cross-model comparisons unreliable, while a lack of diagnostic analyses obscures the internal mechanisms driving transfer efficiency and scaling behaviors. To address this, we introduce EEG-FM-Bench , a unified system for the standardized evaluation of EEG-FMs. The benchmark integrates 14 datasets across 10 paradigms and incorporates diverse experimental settings, including multiple fine-tuning strategies, task organizations, and classifier configurations, supported by tools for gradient and representation analysis. Our experiments and analysis reveal several critical insights: (1) multi-task learning often acts as a useful regularizer that mitigates overfitting in data-scarce EEG contexts, although negative transfer can arise under specific task paradigms; (2) pre-training efficiency is currently limited by gradient conflicts between reconstruction objectives and downstream tasks; (3) under released checkpoints and a matched downstream protocol, model or data scale alone does not fully explain transfer performance, while objective alignment, adaptation compatibility, and EEG-specific design appear to be important factors. This benchmark enables fair comparison and reproducible analysis, providing a step toward fairer comparison and more interpretable analysis of EEG-FMs. Code is available at https://github.com/xw1216/EEG-FM-Bench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a61c3c8c-dca6-4100-bf3d-29937b2a0114Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- MOMENT: A Family of Open Time-series Foundation ModelsMononito Goswami, Konrad Szafer, Arjun Choudhry, Yifu Cai et al.ICML 2024 · 442 citations
- BIOT: Biosignal Transformer for Cross-data Learning in the WildChaoqi Yang, M. Brandon Westover, Jimeng SunNeurIPS 2023 · 345 citations
Related papers
- EmBrace: A Collective Knowledge Fusion Framework Toward Unified EEG Foundation ModelsChenyu Liu, MUYUN JIANG, Pu Wan, Jinxin Pi et al.ICML 2026
- Are EEG Foundation Models Worth It? Comparative Evaluation with Traditional Decoders in Diverse BCI TasksLiuyin Yang, Qiang Sun, Ang Li, Marc M. Van HulleICLR 2026
- NeuroLM: A Universal Multi-task Foundation Model for Bridging the Gap between Language and EEG SignalsWeibang Jiang, Yansen Wang, Bao-Liang Lu, Dongsheng LiICLR 2025
- OSF: On Pre-training and Scaling of Sleep Foundation ModelsZitao Shuai, Zongzhe Xu, David Yang, Wei Wang et al.ICML 2026 · 8 citations
- REVE: A Foundation Model for EEG - Adapting to Any Setup with Large-Scale Pretraining on 25, 000 SubjectsYassine El Ouahidi, Jonathan Lys, Philipp Thölke, Nicolas Farrugia et al.NeurIPS 2025 · 106 citations
