EEG-DLite: Dataset Distillation for Efficient Large EEG Model Training
Yuting Tang, Weibang Jiang, Shanglin Li, Yong Li, Chenyu Liu, Xinliang Zhou, Yi Ding, Cuntai Guan
摘要
Large-scale EEG foundation models have shown strong generalization across a range of downstream tasks, but their training remains resource-intensive due to the volume and variable quality of EEG data. In this work, we introduce EEG-DLite, a data distillation framework that enables more efficient pre-training by selectively removing noisy and redundant samples from large EEG datasets. EEG-DLite begins by encoding EEG segments into compact latent representations using a self-supervised autoencoder, allowing sample selection to be performed efficiently and with reduced sensitivity to noise. Based on these representations, EEG-DLite filters out outliers and minimizes redundancy, resulting in a smaller yet informative subset that retains the diversity essential for effective foundation model training. Through extensive experiments, we demonstrate that training on only 5 percent of a 2,500-hour dataset curated with EEG-DLite yields performance comparable to, and in some cases better than, training on the full dataset across multiple downstream tasks. To our knowledge, this is the first systematic study of pre-training data distillation in the context of EEG foundation models. EEG-DLite provides a scalable and practical path toward more effective and efficient physiological foundation modeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCIWei-Bang Jiang, Li-Ming Zhao, Bao-Liang LuICLR 2024 · 被引用 298 次
- ICE: Inter-instance Contrastive Encoding for Unsupervised Person Re-identificationHao Chen, Benoit Lagadec, François BrémondICCV 2021 · 被引用 258 次
- M3D: Dataset Condensation by Minimizing Maximum Mean DiscrepancyHansong Zhang, Shikun Li, Pengju Wang, Dan Zeng 等AAAI 2024 · 被引用 63 次
- CondTSF: One-line Plugin of Dataset Condensation for Time Series ForecastingJianrong Ding, Zhanyu Liu, Guanjie Zheng, Haiming Jin 等NeurIPS 2024 · 被引用 8 次
- Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep NetworksSiddharth Joshi, Jiayi Ni, Baharan MirzasoleimanICLR 2025
相关 Paper
- REVE: A Foundation Model for EEG - Adapting to Any Setup with Large-Scale Pretraining on 25, 000 SubjectsYassine El Ouahidi, Jonathan Lys, Philipp Thölke, Nicolas Farrugia 等NeurIPS 2025 · 被引用 106 次
- Pre-Training Graph Contrastive Masked Autoencoders are Strong Distillers for EEGXinxu Wei, Kanhao Zhao, Yong Jiao, Hua Xie 等ICML 2025
- LUNA: Efficient and Topology-Agnostic Foundation Model for EEG Signal AnalysisBerkay Döner, Thorir Mar Ingolfsson, Luca Benini, Yawei LiNeurIPS 2025 · 被引用 30 次
- PATCHCODE: Discrete Latent Predictive Learning for EEG Foundation ModelKIEREN YU, Ziyang Liu, Chang Huang, Kaishun WUICML 2026
- Are EEG Foundation Models Worth It? Comparative Evaluation with Traditional Decoders in Diverse BCI TasksLiuyin Yang, Qiang Sun, Ang Li, Marc M. Van HulleICLR 2026
