Finding Shared Decodable Concepts and their Negations in the Brain
Cory Daniel Efird, Alex Murphy, Joel Zylberberg, Alona Fyshe
Abstract
Prior work has offered evidence for functional localization in the brain; different anatomical regions preferentially activate for certain types of visual input. For example, the fusiform face area preferentially activates for visual stimuli that include a face. However, the spectrum of visual semantics is extensive, and only a few semantically-tuned patches of cortex have so far been identified in the human brain. Using a multimodal (natural language and image) neural network architecture (CLIP, Radford et al. (2021), we train a highly accurate contrastive model that maps brain responses during naturalistic image viewing to CLIP embeddings. We then use a novel adaptation of the DBSCAN clustering algorithm to cluster the parameters of these participant-specific contrastive models. This reveals what we call Shared Decodable Concepts (SDCs): clusters in CLIP space that are decodable from common sets of voxels across multiple participants. Examining the images most and least associated with each SDC cluster gives us additional insight into the semantic properties of each SDC. We note SDCs for previously reported visual features (e.g. orientation tuning in early visual cortex) as well as visual semantic concepts such as faces, places and bodies. In cases where our method finds multiple clusters for a visuo-semantic concept, the least associated images allow us to dissociate between confounding factors. For example, we discovered two clusters of food images, one driven by color, the other by shape. We also uncover previously unreported areas with visuo-semantic sensitivity such as regions of extrastriate body area (EBA) tuned for legs/hands and sensitivity to numerosity in right intraparietal sulcus, sensitivity associated with visual perspective (close/far) and more. Thus, our contrastive-learning methodology better characterizes new and existing visuo-semantic representations in the brain by leveraging multimodal neural network representations and a novel adaptation of clustering algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28363c69-d183-450e-856d-d70ef3dbf211Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BrainSCUBA: Fine-Grained Natural Language Captions of Visual Cortex SelectivityAndrew F. Luo, Margaret M. Henderson, Michael J. Tarr, Leila WehbeICLR 2024 · 31 citations
- Brain Dissection: fMRI-trained Networks Reveal Spatial Selectivity in the Processing of Natural ImagesGabriel Sarch, Michael J. Tarr, Katerina Fragkiadaki, Leila WehbeNeurIPS 2023 · 20 citations
Related papers
- CLIP-MSM: A Multi-Semantic Mapping Brain Representation for Human High-Level Visual CortexGuoyuan Yang, Mufan Xue, Ziming Mao, Haofang Zheng et al.AAAI 2025 · 3 citations
- Bridging the Semantic Latent Space between Brain and Machine: Similarity Is All You NeedJiaxuan Chen, Yu Qi, Yueming Wang, Gang PanAAAI 2024 · 13 citations
- Bridging Brains and Concepts: Interpretable Visual Decoding from fMRI with Semantic BottlenecksSara Cammarota, Matteo Ferrante, Nicola ToschiNeurIPS 2025 · 1 citation
- Brain Mapping with Dense Features: Grounding Cortical Semantic Selectivity in Natural Images With Vision TransformersAndrew F. Luo, Jacob Yeung, Rushikesh Zawar, Shaurya Dewan et al.ICLR 2025
- Align2Concept: Language Guided Interpretable Image Recognition by Visual Prototype and Textual Concept AlignmentJiaqi Wang, Pichao Wang, Yi Feng, Huafeng Liu et al.ACM MM 2024 · 1 citation
