Contrast, Attend and Diffuse to Decode High-Resolution Images from Brain Activities
Jingyuan Sun, Mingxiao Li, Zijiao Chen, Yunhao Zhang, Shaonan Wang, Marie-Francine Moens
Abstract
Decoding visual stimuli from neural responses recorded by functional Magnetic Resonance Imaging (fMRI) presents an intriguing intersection between cognitive neuroscience and machine learning, promising advancements in understanding human visual perception and building non-invasive brain-machine interfaces. However, the task is challenging due to the noisy nature of fMRI signals and the intricate pattern of brain visual representations. To mitigate these challenges, we introduce a two-phase fMRI representation learning framework. The first phase pre-trains an fMRI feature learner with a proposed Double-contrastive Mask Auto-encoder to learn denoised representations. The second phase tunes the feature learner to attend to neural activation patterns most informative for visual reconstruction with guidance from an image auto-encoder. The optimized fMRI feature learner then conditions a latent diffusion model to reconstruct image stimuli from brain activities. Experimental results demonstrate our model's superiority in generating high-resolution and semantically accurate images, substantially exceeding previous state-of-the-art methods by 39.34% in the 50-way-top-1 semantic classification accuracy. Our research invites further exploration of the decoding task's potential and contributes to the development of non-invasive brain-machine interfaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- EEG2Video: Towards Decoding Dynamic Visual Perception from EEG SignalsXuan-Hao Liu, Yan-Kai Liu, Yansen Wang, Kan Ren et al.NeurIPS 2024 · 59 citations
- Neuro-Vision to Language: Enhancing Brain Recording-based Visual Reconstruction and Language InteractionGuobin Shen, Dongcheng Zhao, Xiang He, Linghao Feng et al.NeurIPS 2024 · 26 citations
- Exploring Behavior-Relevant and Disentangled Neural Dynamics with Generative Diffusion ModelsYule Wang, Chengrui Li, Weihan Li, Anqi WuNeurIPS 2024 · 13 citations
- NeuralFlix: A Simple While Effective Framework for Semantic Decoding of Videos from Non-invasive Brain RecordingsJingyuan Sun, Mingxiao Li, Marie-Francine MoensAAAI 2025 · 9 citations
- BrainBits: How Much of the Brain are Generative Reconstruction Methods Using?David Mayo, Christopher Wang, Asa Harbin, Abdulrahman Alabdulkareem et al.NeurIPS 2024 · 7 citations
Builds on19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Seeing Beyond the Brain: Conditional Diffusion Model with Sparse Masked Modeling for Vision DecodingZijiao Chen, Jiaxin Qing, Tiange Xiang, Wan Lin Yue et al.CVPR 2023
- MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural DiffusionYizhuo Lu, Changde Du, Qiongyi Zhou, Dianpeng Wang et al.ACM MM 2023 · 48 citations
- Animate Your Thoughts: Reconstruction of Dynamic Natural Vision from Human Brain ActivityYizhuo Lu, Changde Du, Chong Wang, Xuanliu Zhu et al.ICLR 2025
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image ReconstructionXu Zhang, Ruijie Quan, Wenguan Wang, Yi YangICLR 2026
- Visual Decoding and Reconstruction via EEG Embeddings with Guided DiffusionDongyang Li, Chen Wei, Shiying Li, Jiachen Zou et al.NeurIPS 2024 · 164 citations
