Bridging Brains and Concepts: Interpretable Visual Decoding from fMRI with Semantic Bottlenecks
Sara Cammarota, Matteo Ferrante, Nicola Toschi
Abstract
Decoding of visual stimuli from noninvasive neuroimaging techniques such as functional magnetic resonance (fMRI) has advanced rapidly in the last years; yet, most high-performing brain decoding models rely on complicated, non-interpretable latent spaces. In this study we present an interpretable brain decoding framework that inserts a semantic bottleneck into BrainDiffuser, a well established, simple and linear decoding pipeline. We firstly produce a 214 − dimensional binary interpretable space L for images, in which each dimension answers to a specific question about the image (e.g., "Is there a person?", "Is it outdoors?"). A first ridge regression maps voxel activity to this semantic space. Because this mapping is linear, its weight matrix can be visualized as maps of voxel importance for each dimension of L , revealing which cortical regions influence mostly each semantic dimension. A second regression then transforms these concept vectors into CLIP embeddings required to produce the final decoded image, conditioning the Brain-Diffuser model. We found that voxel-wise weight maps for individual questions are highly consistent with canonical category-selective regions in the visual cortex (face, bodies, places, words), simultaneously revealing that activation distributions, not merely location, bear semantic meaning in the brain. Visual brain decoding performance are only slightly lower compared to the original BrainDiffuser metrics (e.g., the CLIP similarity is decreased by ≤ 4% for the four subjects), yet offering substantial gains in interpretability and neuroscientific insights. These results show that our interpretable brain decoding pipeline enables voxel-level analysis of semantic representations in the human brain without sacrificing decoding accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d8cd7c4-e5cc-4056-80d0-aa896c1bdc69Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion PriorsPaul S. Scotti, Atmadeep Banerjee, Jimmie Goode, Stepan Shabalin et al.NeurIPS 2023 · 282 citations
- Versatile Diffusion: Text, Images and Variations All in One Diffusion ModelXingqian Xu, Zhangyang Wang, Eric J. Zhang, Kai Wang et al.ICCV 2023 · 265 citations
Related papers
- MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural DiffusionYizhuo Lu, Changde Du, Qiongyi Zhou, Dianpeng Wang et al.ACM MM 2023 · 48 citations
- CLIP-MSM: A Multi-Semantic Mapping Brain Representation for Human High-Level Visual CortexGuoyuan Yang, Mufan Xue, Ziming Mao, Haofang Zheng et al.AAAI 2025 · 3 citations
- MoRE-Brain: Routed Mixture of Experts for Interpretable and Generalizable Cross-Subject fMRI Visual DecodingYuxiang Wei, Yanteng Zhang, Xi Xiao, Tianyang Wang et al.NeurIPS 2025 · 15 citations
- Towards Interpretable Visual Decoding with Attention to Brain RepresentationsPinyuan Feng, Hossein Adeli, Wenxuan Guo, Fan Cheng et al.ICLR 2026 · 1 citation
- Seeing Beyond the Brain: Conditional Diffusion Model with Sparse Masked Modeling for Vision DecodingZijiao Chen, Jiaxin Qing, Tiange Xiang, Wan Lin Yue et al.CVPR 2023
