EEG-DLite: Dataset Distillation for Efficient Large EEG Model Training
Yuting Tang, Weibang Jiang, Shanglin Li, Yong Li, Chenyu Liu, Xinliang Zhou, Yi Ding, Cuntai Guan
Abstract
Large-scale EEG foundation models have shown strong generalization across a range of downstream tasks, but their training remains resource-intensive due to the volume and variable quality of EEG data. In this work, we introduce EEG-DLite, a data distillation framework that enables more efficient pre-training by selectively removing noisy and redundant samples from large EEG datasets. EEG-DLite begins by encoding EEG segments into compact latent representations using a self-supervised autoencoder, allowing sample selection to be performed efficiently and with reduced sensitivity to noise. Based on these representations, EEG-DLite filters out outliers and minimizes redundancy, resulting in a smaller yet informative subset that retains the diversity essential for effective foundation model training. Through extensive experiments, we demonstrate that training on only 5 percent of a 2,500-hour dataset curated with EEG-DLite yields performance comparable to, and in some cases better than, training on the full dataset across multiple downstream tasks. To our knowledge, this is the first systematic study of pre-training data distillation in the context of EEG foundation models. EEG-DLite provides a scalable and practical path toward more effective and efficient physiological foundation modeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1dc5fe94-c005-4acd-8542-cfdb4f1d1a33Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCIWei-Bang Jiang, Li-Ming Zhao, Bao-Liang LuICLR 2024 · 298 citations
- ICE: Inter-instance Contrastive Encoding for Unsupervised Person Re-identificationHao Chen, Benoit Lagadec, François BrémondICCV 2021 · 258 citations
- M3D: Dataset Condensation by Minimizing Maximum Mean DiscrepancyHansong Zhang, Shikun Li, Pengju Wang, Dan Zeng et al.AAAI 2024 · 63 citations
- CondTSF: One-line Plugin of Dataset Condensation for Time Series ForecastingJianrong Ding, Zhanyu Liu, Guanjie Zheng, Haiming Jin et al.NeurIPS 2024 · 8 citations
- Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep NetworksSiddharth Joshi, Jiayi Ni, Baharan MirzasoleimanICLR 2025
Related papers
- REVE: A Foundation Model for EEG - Adapting to Any Setup with Large-Scale Pretraining on 25, 000 SubjectsYassine El Ouahidi, Jonathan Lys, Philipp Thölke, Nicolas Farrugia et al.NeurIPS 2025 · 106 citations
- Pre-Training Graph Contrastive Masked Autoencoders are Strong Distillers for EEGXinxu Wei, Kanhao Zhao, Yong Jiao, Hua Xie et al.ICML 2025
- LUNA: Efficient and Topology-Agnostic Foundation Model for EEG Signal AnalysisBerkay Döner, Thorir Mar Ingolfsson, Luca Benini, Yawei LiNeurIPS 2025 · 30 citations
- PATCHCODE: Discrete Latent Predictive Learning for EEG Foundation ModelKIEREN YU, Ziyang Liu, Chang Huang, Kaishun WUICML 2026
- Are EEG Foundation Models Worth It? Comparative Evaluation with Traditional Decoders in Diverse BCI TasksLiuyin Yang, Qiang Sun, Ang Li, Marc M. Van HulleICLR 2026
