A Multi-view Spectral-Spatial-Temporal Masked Autoencoder for Decoding Emotions with Self-supervised Learning
Rui Li, Yiting Wang, Wei-Long Zheng, Bao-Liang Lu
Abstract
Affective Brain-computer Interface has achieved considerable advances that researchers can successfully interpret labeled and flawless EEG data collected in laboratory settings. However, the annotation of EEG data is time-consuming and requires a vast workforce which limits the application in practical scenarios. Furthermore, daily collected EEG data may be partially damaged since EEG signals are sensitive to noise. In this paper, we propose a Multi-view Spectral-Spatial-Temporal Masked Autoencoder (MV-SSTMA) with self-supervised learning to tackle these challenges towards daily applications. The MV-SSTMA is based on a multi-view CNN-Transformer hybrid structure, interpreting the emotion-related knowledge of EEG signals from spectral, spatial, and temporal perspectives. Our model consists of three stages: 1) In the generalized pre-training stage, channels of unlabeled EEG data from all subjects are randomly masked and later reconstructed to learn the generic representations from EEG data; 2) In the personalized calibration stage, only few labeled data from a specific subject are used to calibrate the model; 3) In the personal test stage, our model can decode personal emotions from the sound EEG data as well as damaged ones with missing channels. Extensive experiments on two open emotional EEG datasets demonstrate that our proposed model achieves state-of-the-art performance on emotion recognition. In addition, under the abnormal circumstance of missing channels, the proposed model can still effectively recognize emotions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 044f3373-cb39-4b33-bc8f-4004e44fede0Cited by top-tier papers5
- Learning Topology-Agnostic EEG Representations with Geometry-Aware ModelingKe Yi, Yansen Wang, Kan Ren, Dongsheng LiNeurIPS 2023 · 99 citations
- REmoNet: Reducing Emotional Label Noise via Multi-regularized Self-supervisionWei-Bang Jiang, Yu-Ting Lan, Bao-Liang LuACM MM 2024 · 3 citations
- Brain-Inspired fMRI-to-Text Decoding via Incremental and Wrap-Up Language ModelingWentao Lu, Dong Nie, Pengcheng Xue, Zheng Cui et al.NeurIPS 2025 · 3 citations
- EEG-SCMM: Soft Contrastive Masked Modeling for Cross-Corpus EEG-Based Emotion RecognitionQile Liu, Weishan Ye, Lingli Zhang, Zhen LiangACM MM 2025 · 1 citation
- Enhancing EEG-to-Text Decoding through Transferable Representations from Pre-trained Contrastive EEG-Text Masked AutoencoderJiaqi Wang, Zhenxi Song, Zhengyu Ma, Xipeng Qiu et al.ACL 2024
Builds on4
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- SST-EmotionNet: Spatial-Spectral-Temporal based Attention 3D Dense Network for EEG Emotion RecognitionZiyu Jia, Youfang Lin, Xiyang Cai, Haobin Chen et al.ACM MM 2020 · 166 citations
- A Multi-Domain Adaptive Graph Convolutional Network for EEG-based Emotion RecognitionRui Li, Yiting Wang, Bao-Liang LuACM MM 2021 · 70 citations
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li et al.CVPR 2022
Related papers
- Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity RecognitionShenghuan Miao, Ling Chen, Rong HuUbiComp 2024 · 23 citations
- Pretraining Large Brain Language Model for Active BCI: Silent SpeechJinzhao Zhou, Zehong Cao, Yiqun Duan, Connor Barkley et al.ACM MM 2025 · 2 citations
- WSEL: EEG Feature Selection with Weighted Self-expression Learning for Incomplete Multi-dimensional Emotion RecognitionXueyuan Xu, Li Zhuo, Jinxin Lu, Xia WuACM MM 2024 · 2 citations
- DMMR: Cross-Subject Domain Generalization for EEG-Based Emotion Recognition via Denoising Mixed Mutual ReconstructionYiming Wang, Bin Zhang, Yujiao TangAAAI 2024 · 51 citations
- SEBSFormer: A Spectral-Enhanced Bi-Stream Transformer for Robust EEG DecodingLin Zhang, Shikui Tu, Lei XuAAAI 2026
