Visual Decoding and Reconstruction via EEG Embeddings with Guided Diffusion
Dongyang Li, Chen Wei, Shiying Li, Jiachen Zou, Quanying Liu
Abstract
How to decode human vision through neural signals has attracted a long-standing interest in neuroscience and machine learning. Modern contrastive learning and generative models improved the performance of visual decoding and reconstruction based on functional Magnetic Resonance Imaging (fMRI). However, the high cost and low temporal resolution of fMRI limit their applications in brain-computer interfaces (BCIs), prompting a high need for visual decoding based on electroencephalography (EEG). In this study, we present an end-to-end EEG-based visual reconstruction zero-shot framework, consisting of a tailored brain encoder, called the Adaptive Thinking Mapper (ATM), which projects neural signals from different sources into the shared subspace as the clip embedding, and a two-stage multi-pipe EEG-to-image generation strategy. In stage one, EEG is embedded to align the high-level clip embedding, and then the prior diffusion model refines EEG embedding into image priors. A blurry image also decoded from EEG for maintaining the low-level feature. In stage two, we input both the high-level clip embedding, the blurry image and caption from EEG latent to a pre-trained diffusion model. Furthermore, we analyzed the impacts of different time windows and brain regions on decoding and reconstruction. The versatility of our framework is demonstrated in the magnetoencephalogram (MEG) data modality. The experimental results indicate that our EEG-based visual zero-shot framework achieves SOTA performance in classification, retrieval and reconstruction, highlighting the portability, low cost, and high temporal resolution of EEG, enabling a wide range of BCI applications. Our code is available at https://github.com/ncclab-sustech/EEG_Image_decode.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c52ef225-321f-4052-a21e-e4e0fe81c9c9Cited by top-tier papers32
- CognitionCapturer: Decoding Visual Stimuli from Human EEG Signal with Multimodal InformationKaifan Zhang, Lihuo He, Xin Jiang, Wen Lu et al.AAAI 2025 · 34 citations
- Meta-Learning an In-Context Transformer Model of Human Higher Visual CortexMuquan Yu, Mu Nan, Hossein Adeli, Jacob S. Prince et al.NeurIPS 2025 · 5 citations
- EEGMirror: Leveraging EEG Data in the Wild Via Montage-Agnostic Self-Supervision for EEG to Video DecodingXuan-Hao Liu, Bao-Liang Lu, Wei-Long ZhengICCV 2025 · 5 citations
- SATTC: Structure-Aware Label-Free Test-Time Calibration for Cross-Subject EEG-to-Image RetrievalQunjie Huang, Weina ZhuCVPR 2026 · 4 citations
- Autoregressive Visual Decoding from EEG SignalsSicheng Dai, Hongwang Xiao, Shan Yu, Qiwei YeICLR 2026 · 3 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- Leveraging Visual Blur Perception Characteristics for EEG DecodingWenchao Liu, Hongwei Li, Zhouyang Xu, Lin Ma et al.AAAI 2026
- Decoding Natural Images from EEG for Object RecognitionYonghao Song, Bingchuan Liu, Xiang Li, Nanlin Shi et al.ICLR 2024 · 135 citations
- NEED: Cross-Subject and Cross-Task Generalization for Video and Image Reconstruction from EEG SignalsShuai Huang, Huan Luo, Haodong Jing, Qixian Zhang et al.NeurIPS 2025 · 17 citations
- EVOKE: Efficient and High-Fidelity EEG-to-Video Reconstruction via Decoupling Implicit Neural RepresentationHaodong Jing, Panqi Yang, Dongyao Jiang, Zhipeng Liu et al.AAAI 2026 · 1 citation
- MB2C: Multimodal Bidirectional Cycle Consistency for Learning Robust Visual Neural RepresentationsYayun Wei, Lei Cao, Hao Li, Yilin DongACM MM 2024 · 21 citations
