EEGMirror: Leveraging EEG Data in the Wild Via Montage-Agnostic Self-Supervision for EEG to Video Decoding
Xuan-Hao Liu, Bao-Liang Lu, Wei-Long Zheng
Abstract
Generating high fidelity video from brain activity is an important milestone in brain decoding research. Previous works were mostly based on functional Magnetic Resonance Imaging (fMRI), whose low temporal resolution confines the ability of faithfully reflecting rapid brain activity, motivating us to turn to high temporal resolution brain signals like electroencephalography (EEG). However, EEGto-video is challenging due to the complexity and nonstationarity of EEG signals and the scarcity of annotated data. Addressing these issues, we present EEGMirror. Firstly, we adopt neural quantization to convert nonstationary signals into robust discrete representation. Afterwards, a masked self-supervision method with montage-agnostic position embedding (MAPE) is introduced to acquire an effective EEG encoder. By MAPE, our model can flexibly leverage different EEG datasets with various montages (number and position of channels), mitigating the lack of wellannotated data. Next, multimodal contrastive learning is applied to align the brain modality with dynamic changes and semantic information. Lastly, a fine-tuned inflated Stable Diffusion model is adopted to reconstruct video stimuli guided by visual and semantic information decoded from EEG signals. We show that EEGMirror outperforms the state-of-the-art performance in both semantic (82.1% vs 79.8%) and pixel (0.261 vs 0.256) levels. An exhaustive ablation study is also conducted to analyze our framework. Code: https://github.com/XuanhaoLiu/EEGMirror.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d4f835b-c505-47ae-8a7c-2acf325e3d5cCited by top-tier papers1
Ask how each one uses itBuilds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- EEG2Video: Towards Decoding Dynamic Visual Perception from EEG SignalsXuan-Hao Liu, Yan-Kai Liu, Yansen Wang, Kan Ren et al.NeurIPS 2024 · 59 citations
- EVOKE: Efficient and High-Fidelity EEG-to-Video Reconstruction via Decoupling Implicit Neural RepresentationHaodong Jing, Panqi Yang, Dongyao Jiang, Zhipeng Liu et al.AAAI 2026 · 1 citation
- Visual Decoding and Reconstruction via EEG Embeddings with Guided DiffusionDongyang Li, Chen Wei, Shiying Li, Jiachen Zou et al.NeurIPS 2024 · 164 citations
- Decoding Natural Images from EEG for Object RecognitionYonghao Song, Bingchuan Liu, Xiang Li, Nanlin Shi et al.ICLR 2024 · 135 citations
- MINDEV: Multi-modal Integrated Diffusion Framework for Video Reconstruction from EEG SignalsShuai Huang, Yongxiong Wang, Huan Luo, Haodong Jing et al.ACM MM 2025 · 1 citation
