Alleviating the Semantic Gap for Generalized fMRI-to-Image Reconstruction
Tao Fang, Qian Zheng, Gang Pan
Abstract
Although existing fMRI-to-image reconstruction methods could predict high-quality images, they do not explicitly consider the semantic gap between training and testing data, resulting in reconstruction with unstable and uncertain semantics. This paper addresses the problem of generalized fMRI-to-image reconstruction by explicitly alleviates the semantic gap. Specifically, we leverage the pre-trained CLIP model to map the training data to a compact feature representation, which essentially extends the sparse semantics of training data to dense ones, thus alleviating the semantic gap of the instances nearby known concepts (i.e., inside the training super-classes). Inspired by the robust low-level representation in fMRI data, which could help alleviate the semantic gap for instances that far from the known concepts (i.e., outside the training super-classes), we leverage structural information as a general cue to guide image reconstruction. Further, we quantify the semantic uncertainty based on probability density estimation and achieve G eneralized fMRI-to-image reconstruction by adaptively integrating E xpanded S emantics and S tructural information ( GESS ) within a diffusion process. Experimental results demonstrate that the proposed GESS model outperforms state-of-the-art methods, and we propose a generalized scenario split strategy to evaluate the advantage of GESS in closing the semantic gap. Our codes are available at https://github.com/duolala1/GESS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 567c205f-5bdd-4205-a269-1baf0643e158Cited by top-tier papers14
- Visual Decoding and Reconstruction via EEG Embeddings with Guided DiffusionDongyang Li, Chen Wei, Shiying Li, Jiachen Zou et al.NeurIPS 2024 · 164 citations
- NEED: Cross-Subject and Cross-Task Generalization for Video and Image Reconstruction from EEG SignalsShuai Huang, Huan Luo, Haodong Jing, Qixian Zhang et al.NeurIPS 2025 · 17 citations
- Resisting Stochastic Risks in Diffusion Planners with the Trajectory Aggregation TreeLang Feng, Pengjie Gu, Bo An, Gang PanICML 2024 · 15 citations
- Bridging the Semantic Latent Space between Brain and Machine: Similarity Is All You NeedJiaxuan Chen, Yu Qi, Yueming Wang, Gang PanAAAI 2024 · 13 citations
- ZEBRA: Towards Zero-Shot Cross-Subject Generalization for Universal Brain Visual DecodingHaonan Wang, Jingyu Lu, Hongrui Li, Xiaomeng LiNeurIPS 2025 · 8 citations
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural DiffusionYizhuo Lu, Changde Du, Qiongyi Zhou, Dianpeng Wang et al.ACM MM 2023 · 48 citations
- NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video ReconstructionZixuan Gong, Guangyin Bao, Qi Zhang, Zhongwei Wan et al.NeurIPS 2024 · 39 citations
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image ReconstructionXu Zhang, Ruijie Quan, Wenguan Wang, Yi YangICLR 2026
- Bridging the Gap Between Brain and Machine in Interpreting Visual Semantics: Towards Self-Adaptive Brain-to-Text DecodingJiaxuan Chen, Yu Qi, Yueming Wang, Gang PanICCV 2025 · 1 citation
- Mind Reader: Reconstructing complex images from brain activitiesSikun Lin, Thomas Sprague, Ambuj K. SinghNeurIPS 2022 · 155 citations
