BrainSCUBA: Fine-Grained Natural Language Captions of Visual Cortex Selectivity
Andrew F. Luo, Margaret M. Henderson, Michael J. Tarr, Leila Wehbe
Abstract
Understanding the functional organization of higher visual cortex is a central focus in neuroscience. Past studies have primarily mapped the visual and semantic selectivity of neural populations using hand-selected stimuli, which may potentially bias results towards pre-existing hypotheses of visual cortex functionality. Moving beyond conventional approaches, we introduce a data-driven method that generates natural language descriptions for images predicted to maximally activate individual voxels of interest. Our method -Semantic Captioning Using Brain Alignments ("BrainSCUBA") -builds upon the rich embedding space learned by a contrastive vision-language model and utilizes a pre-trained large language model to generate interpretable captions. We validate our method through fine-grained voxel-level captioning across higher-order visual regions. We further perform text-conditioned image synthesis with the captions, and show that our images are semantically coherent and yield high predicted activations. Finally, to demonstrate how our method enables scientific discovery, we perform exploratory investigations on the distribution of "person" representations in the brain, and discover fine-grained semantic selectivity in body-selective areas. Unlike earlier studies that decode text, our method derives voxel-wise captions of semantic selectivity. Our results show that BrainSCUBA is a promising means for understanding functional preferences in the brain, and provides motivation for further hypothesis-driven investigation of visual cortex. https://www.cs.cmu.edu/ ˜afluo/BrainSCUBA
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e245f04-78b0-4bab-865a-e05ee76cd9deCited by top-tier papers21
- MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of DataPaul S. Scotti, Mihir Tripathy, Cesar Torrico, Reese Kneeland et al.ICML 2024 · 117 citations
- Neuro-Vision to Language: Enhancing Brain Recording-based Visual Reconstruction and Language InteractionGuobin Shen, Dongcheng Zhao, Xiang He, Linghao Feng et al.NeurIPS 2024 · 26 citations
- Transformer brain encoders explain human high-level visual responsesHossein Adeli, Minni Sun, Nikolaus KriegeskorteNeurIPS 2025 · 14 citations
- Brain Decodes Deep NetsHuzheng Yang, James C. Gee, Jianbo ShiCVPR 2024 · 10 citations
- LaVCa: LLM-assisted Visual Cortex CaptioningTakuya Matsuyama, Shinji Nishimoto, Yu TakagiICLR 2026 · 8 citations
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal et al.ICLR 2022 · 503 citations
- ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic ArithmeticYoad Tewel, Yoav Shalev, Idan Schwartz, Lior WolfCVPR 2022 · 129 citations
- Brain Diffusion for Visual Exploration: Cortical Discovery using Large Scale Generative ModelsAndrew F. Luo, Margaret M. Henderson, Leila Wehbe, Michael J. TarrNeurIPS 2023 · 54 citations
Related papers
- Brain Mapping with Dense Features: Grounding Cortical Semantic Selectivity in Natural Images With Vision TransformersAndrew F. Luo, Jacob Yeung, Rushikesh Zawar, Shaurya Dewan et al.ICLR 2025
- MindSimulator: Exploring Brain Concept Localization via Synthetic fMRIGuangyin Bao, Qi Zhang, Zixuan Gong, Zhuojia Wu et al.ICLR 2025
- BrainLMM: A Label-Free Framework for Mapping Multi-Semantic Representation in the Human Visual CortexTan Gao, Mufan Xue, Haofang Zheng, Shuo Lv et al.AAAI 2026
- In Silico Mapping of Visual Categorical Selectivity Across the Whole BrainEthan Hwang, Hossein Adeli, Wenxuan Guo, Andrew F. Luo et al.NeurIPS 2025 · 7 citations
- Finding Shared Decodable Concepts and their Negations in the BrainCory Daniel Efird, Alex Murphy, Joel Zylberberg, Alona FysheICLR 2025
