D-FOSA: Dual-Diffusion Guided EEG-to-Image Reconstruction with Frequency-Oriented Semantic Alignment
Chenglong Yu, Shuai Shen, Xiangsheng Li, Yang Li
Abstract
Reconstructing visual semantics from Electroencephalography (EEG) signals enables a deeper understanding of human visual cognition and supports next-generation brain–computer interface (BCI) applications.Despite notable advances in recent years, most existing EEG encoders still struggle to capture the frequency-specific neural dynamics that reflect perceptual and cognitive rhythms. Moreover, the cross-modal alignment between EEG and visual content remains insufficiently tackled, leading to limited semantic consistency and visual fidelity. To address these issues, we propose D-FOSA, a unified dual-diffusion guided framework with frequency-oriented semantic alignment, which strengthens the frequency-aware EEG representation for more semantically aligned image reconstruction.Specifically, we design a Frequency-Spatio-Temporal Dynamics Encoder (FSTDE) based on the Frequency-Oriented Mamba (FOMamba) to explicitly model oscillatory patterns and long-range dependencies in EEG signals. The extracted features are then pulled into the CLIP-aligned visual semantic space via contrastive learning.Meanwhile, a Dual Diffusion Latent Generator (DDLG) with bidirectional EEG–image conditioning is designed to enforce cross-modal alignment and promote cycle-consistent generation.Extensive experiments on four challenging datasets demonstrate that our proposed D-FOSA significantly outperforms existing methods in both retrieval and reconstruction tasks. Particularly, our D-FOSA surpasses the contemporary MB2C method by over 20 FID in the reconstruction task on THINGS-EEG, indicating a substantial improvement in perceptual fidelity. The source code is in the supplementary material.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22a2a442-31d2-466a-856f-b4383b863248Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- MB2C: Multimodal Bidirectional Cycle Consistency for Learning Robust Visual Neural RepresentationsYayun Wei, Lei Cao, Hao Li, Yilin DongACM MM 2024 · 21 citations
- Visual Decoding and Reconstruction via EEG Embeddings with Guided DiffusionDongyang Li, Chen Wei, Shiying Li, Jiachen Zou et al.NeurIPS 2024 · 164 citations
- MINDEV: Multi-modal Integrated Diffusion Framework for Video Reconstruction from EEG SignalsShuai Huang, Yongxiong Wang, Huan Luo, Haodong Jing et al.ACM MM 2025 · 1 citation
- Contrast, Attend and Diffuse to Decode High-Resolution Images from Brain ActivitiesJingyuan Sun, Mingxiao Li, Zijiao Chen, Yunhao Zhang et al.NeurIPS 2023 · 57 citations
- Leveraging Visual Blur Perception Characteristics for EEG DecodingWenchao Liu, Hongwei Li, Zhouyang Xu, Lin Ma et al.AAAI 2026
