Revealing Vision-Language Integration in the Brain with Multimodal Networks
Vighnesh Subramaniam, Colin Conwell, Christopher Wang, Gabriel Kreiman, Boris Katz, Ignacio Cases, Andrei Barbu
Abstract
We use (multi)modal deep neural networks (DNNs) to probe for sites of multimodal integration in the human brain by predicting stereoen-cephalography (SEEG) recordings taken while human subjects watched movies. We operationalize sites of multimodal integration as regions where a multimodal vision-language model predicts recordings better than unimodal language, unimodal vision, or linearly-integrated language-vision models. Our target DNN models span different architectures (e.g., convolutional networks and transformers) and multimodal training techniques (e.g., cross-attention and contrastive learning). As a key enabling step, we first demonstrate that trained vision and language models systematically outperform their randomly initialized counterparts in their ability to predict SEEG signals. We then compare unimodal and multimodal models against one another. Because our target DNN models often have different architectures, number of parameters, and training sets (possibly obscuring those differences attributable to integration), we carry out a controlled comparison of two models (SLIP and SimCLR), which keep all of these attributes the same aside from input modality. Using this approach, we identify a sizable number of neural sites (on average 141 out of 1090 total sites or 12.94%) and brain regions where multimodal integration seems to occur. Additionally, we find that among the variants of multimodal training techniques we assess, CLIP-style training is the best suited for downstream prediction of the neural activity in these sites.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7fd18dd6-3e5c-4b4f-8170-5fa55a088856Cited by top-tier papers6
- Quantifying Task-relevant Similarities in Representations Using Decision Variable CorrelationsYu Qian, Wilson S. Geisler, Xue-Xin WeiNeurIPS 2025 · 1 citation
- Each Complexity Deserves a Pruning PolicyHanshi Wang, Yuhao Xu, Zekun Xu, Jin Gao et al.NeurIPS 2025 · 1 citation
- Multi-modal brain encoding models for multi-modal stimuliSubba Reddy Oota, Khushbu Pahwa, Mounika Marreddy, Maneesh Kumar Singh et al.ICLR 2025
- Jacobian-Based Interpretation of Nonlinear Neural Encoding ModelXiaohui Gao, Haoran Yang, Yue Cheng, Mengfei Zuo et al.NeurIPS 2025
- Dual Coding Theory in Action: Language-Assisted Human Pose Estimation in VideosSifan Wu, Haipeng Chen, Yingda Lyu, Shaojing Fan et al.AAAI 2026
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon et al.CVPR 2022 · 483 citations
Related papers
- SIM: Surface-based fMRI Analysis for Inter-Subject Multimodal Decoding from Movie-Watching ExperimentsSimon Dahan, Gabriel Bénédict, Logan Zane John Williams, Yourong Guo et al.ICLR 2025
- Brain encoding models based on multimodal transformers can transfer across language and visionJerry Tang, Meng Du, Vy A. Vo, Vasudev Lal et al.NeurIPS 2023 · 76 citations
- How can embedding models bind concepts?Arnas Uselis, Darina Koishigarina, Seong Joon OhICML 2026
- i-Code: An Integrative and Composable Multimodal Learning FrameworkZiyi Yang, Yuwei Fang, Chenguang Zhu, Reid Pryzant et al.AAAI 2023 · 53 citations
- Finding Shared Decodable Concepts and their Negations in the BrainCory Daniel Efird, Alex Murphy, Joel Zylberberg, Alona FysheICLR 2025
