Is Limited Participant Diversity Impeding EEG-based Machine Learning?
Philipp Bomatter, Henry Gouk
Abstract
The application of machine learning (ML) to electroencephalography (EEG) has great potential to advance both neuroscientific research and clinical applications. However, the generalisability and robustness of EEG-based ML models often hinge on the amount and diversity of training data. It is common practice to split EEG recordings into small segments, thereby increasing the number of samples substantially compared to the number of individual recordings or participants. We conceptualise this as a multi-level data generation process and investigate the scaling behaviour of model performance with respect to the overall sample size and the participant diversity through large-scale empirical studies. We then use the same framework to investigate the effectiveness of different ML strategies designed to address limited data problems: data augmentations and self-supervised learning. Our findings show that model performance scaling can be severely constrained by participant distribution shifts and provide actionable guidance for data collection and ML research. The code for our experiments is publicly available online. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db7d702c-de15-4f80-b60a-bd1732b60982Builds on8
- BIOT: Biosignal Transformer for Cross-data Learning in the WildChaoqi Yang, M. Brandon Westover, Jimeng SunNeurIPS 2023 · 345 citations
- Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCIWei-Bang Jiang, Li-Ming Zhao, Bao-Liang LuICLR 2024 · 298 citations
- EEGPT: Pretrained Transformer for Universal and Reliable Representation of EEG SignalsGuangyu Wang, Wenchao Liu, Yuhong He, Cong Xu et al.NeurIPS 2024 · 267 citations
- MAtt: A Manifold Attention Network for EEG DecodingYue-Ting Pan, Jing-Lun Chou, Chun-Shu WeiNeurIPS 2022 · 105 citations
- CADDA: Class-wise Automatic Differentiable Data Augmentation for EEG SignalsCédric Rommel, Thomas Moreau, Joseph Paillard, Alexandre GramfortICLR 2022 · 52 citations
Related papers
- Quantifying the Generalization Gap in Seizure Detection: A Large-Scale Empirical Benchmark via the SzCORE ChallengeJonathan Dan, Amirhossein Shahbazinia, Christodoulos Kechris, David AtienzaICML 2026 · 4 citations
- EEG Agent: A Unified Framework for Automated EEG Analysis Using Large Language ModelsSha Zhao, Mingyi Peng, Haiteng Jiang, Tao Li et al.AAAI 2026
- The Brain's Bitter Lesson: Scaling Speech Decoding With Self-Supervised LearningDulhan Jayalath, Gilad Landau, Brendan Shillingford, Mark W. Woolrich et al.ICML 2025
- A foundation model with multi-variate parallel attention to generate neuronal activityFrancesco S. Carzaniga, Michael Hersche, Abu Sebastian, Kaspar Schindler et al.ICLR 2026 · 6 citations
- CodeBrain: Bridging Decoupled Tokenizer and Multi-Scale Architecture for EEG Foundation ModelJingying Ma, Feng Wu, Qika Lin, Yucheng Xing et al.ICLR 2026 · 25 citations
