CLIP-MSM: A Multi-Semantic Mapping Brain Representation for Human High-Level Visual Cortex
Guoyuan Yang, Mufan Xue, Ziming Mao, Haofang Zheng, Jia Xu, Dabin Sheng, Ruotian Sun, Ruoqi Yang, Xuesong Li
Abstract
Prior work employing deep neural networks (DNNs) with explainable techniques has identified human visual cortical selective representation to specific categories. However, constructing high-performing encoding models that accurately capture brain responses to coexisting multi-semantics remains elusive. Here, we used CLIP models combined with CLIP Dissection to establish a multi-semantic mapping framework (CLIP-MSM) for hypothesis-free analysis in human high-level visual cortex. First, we utilize CLIP models to construct voxel-wise encoding models for predicting visual cortical responses to natural scene images. Then, we apply CLIP Dissection and normalize the semantic mapping score to achieve the mapping of single brain voxels to multiple semantics. Our findings indicate that CLIP Dissection applied to DNNs modeling the human high-level visual cortex demonstrates better interpretability accuracy compared to Network Dissection. In addition, to demonstrate how our method enables fine-grained discovery in hypothesis-free analysis, we quantify the accuracy between CLIP-MSM’s reconstructed brain activation in response to categories of faces, bodies, places, words and food, and the ground truth of brain activation. We demonstrate that CLIP-MSM provides more accurate predictions of visual responses compared to CLIP Dissection. Our results have been validated using two large natural image datasets: the Natural Scenes Dataset (NSD) and the Natural Object Dataset (NOD).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7a9b8c1-33e3-4e18-b243-56c5e6690c68Cited by top-tier papers2
- BrainLMM: A Label-Free Framework for Mapping Multi-Semantic Representation in the Human Visual CortexTan Gao, Mufan Xue, Haofang Zheng, Shuo Lv et al.AAAI 2026
- SAEs-BrainMap: Unveiling the Emergence of Specialized Concepts in Deep Models via Brain AlignmentZiming Mao, Jia Xu, Wenxuan Pan, Mufan Xue et al.ICML 2026
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- Toward a realistic model of speech processing in the brain with self-supervised learningJuliette Millet, Charlotte Caucheteux, Pierre Orhan, Yves Boubenec et al.NeurIPS 2022 · 164 citations
- Brain Diffusion for Visual Exploration: Cortical Discovery using Large Scale Generative ModelsAndrew F. Luo, Margaret M. Henderson, Leila Wehbe, Michael J. TarrNeurIPS 2023 · 54 citations
- BrainSCUBA: Fine-Grained Natural Language Captions of Visual Cortex SelectivityAndrew F. Luo, Margaret M. Henderson, Michael J. Tarr, Leila WehbeICLR 2024 · 31 citations
Related papers
- Finding Shared Decodable Concepts and their Negations in the BrainCory Daniel Efird, Alex Murphy, Joel Zylberberg, Alona FysheICLR 2025
- Bridging Brains and Concepts: Interpretable Visual Decoding from fMRI with Semantic BottlenecksSara Cammarota, Matteo Ferrante, Nicola ToschiNeurIPS 2025 · 1 citation
- CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision NetworksTuomas P. Oikarinen, Tsui-Wei WengICLR 2023 · 9 citations
- Brain Dissection: fMRI-trained Networks Reveal Spatial Selectivity in the Processing of Natural ImagesGabriel Sarch, Michael J. Tarr, Katerina Fragkiadaki, Leila WehbeNeurIPS 2023 · 20 citations
- Brain Mapping with Dense Features: Grounding Cortical Semantic Selectivity in Natural Images With Vision TransformersAndrew F. Luo, Jacob Yeung, Rushikesh Zawar, Shaurya Dewan et al.ICLR 2025
