Brain encoding models based on multimodal transformers can transfer across language and vision
Jerry Tang, Meng Du, Vy A. Vo, Vasudev Lal, Alexander Huth
Abstract
Encoding models have been used to assess how the human brain represents concepts in language and vision. While language and vision rely on similar concept representations, current encoding models are typically trained and tested on brain responses to each modality in isolation. Recent advances in multimodal pretraining have produced transformers that can extract aligned representations of concepts in language and vision. In this work, we used representations from multimodal transformers to train encoding models that can transfer across fMRI responses to stories and movies. We found that encoding models trained on brain responses to one modality can successfully predict brain responses to the other modality, particularly in cortical regions that represent conceptual meaning. Further analysis of these encoding models revealed shared semantic dimensions that underlie concept representations in language and vision. Comparing encoding models trained using representations from multimodal and unimodal transformers, we found that multimodal transformers learn more aligned representations of concepts in language and vision. Our results demonstrate how multimodal transformers can provide insights into the brain's capacity for multimodal processing. Encoding models predict brain responses from quantitative features of the stimuli that elicited them [1] . In recent years, fitting encoding models to data from functional magnetic resonance imaging (fMRI) experiments has become a powerful approach for understanding information processing in the brain. While encoding models are usually trained and tested on brain responses to a single stimulus modality, such as language [2-8] or vision [9] [10] [11] [12] [13] [14] , the human brain is remarkable in its ability to integrate information across multiple modalities. There is growing evidence that this capacity for multimodal processing is supported by aligned cortical representations of the same concepts in different modalities-for instance, hearing the sentence "a dog chases a cat" and seeing a dog chasing a cat may elicit similar patterns of brain activity [15] [16] [17] [18] [19] [20] . In this work, we investigated the alignment between language and visual representations in the brain by training encoding models on fMRI responses to one modality and testing them on fMRI responses to the other modality. Encoding models that successfully transfer across modalities can provide insights into how the two modalities are related [19] . Although previous work has compared language and vision encoding models, human annotations were required to map language and visual stimuli into a shared semantic space [19] . To our knowledge, cross-modality transfer has yet to be demonstrated using encoding models trained on stimulus-computable features that capture the rich connections between language and vision. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- TRIBE: TRImodal Brain Encoder for whole-brain fMRI response predictionStéphane d'Ascoli, Jérémy Rapin, Yohann Benchetrit, Hubert Banville et al.ICLR 2026 · 28 citations
- Can Transformers Smell Like Humans?Farzaneh Taleb, Miguel Vasco, Antônio H. Ribeiro, Mårten Björkman et al.NeurIPS 2024 · 11 citations
- SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation LearningWeijian Mai, Jiamin Wu, Yu Zhu, Zhouheng Yao et al.NeurIPS 2025 · 11 citations
- In Silico Mapping of Visual Categorical Selectivity Across the Whole BrainEthan Hwang, Hossein Adeli, Wenxuan Guo, Andrew F. Luo et al.NeurIPS 2025 · 7 citations
- Disentangling Superpositions: Interpretable Brain Encoding Model with Sparse Concept AtomsAlicia Zeng, Jack GallantNeurIPS 2025 · 5 citations
Builds on4
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- BridgeTower: Building Bridges between Encoders in Vision-Language Representation LearningXiao Xu, Chenfei Wu, Shachar Rosenman, Vasudev Lal et al.AAAI 2023 · 99 citations
Related papers
- Multi-modal brain encoding models for multi-modal stimuliSubba Reddy Oota, Khushbu Pahwa, Mounika Marreddy, Maneesh Kumar Singh et al.ICLR 2025
- Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)Subba Reddy Oota, Akshett Rai Jindal, Ishani Mondal, Khushbu Pahwa et al.ICLR 2025
- Low-dimensional Structure in the Space of Language Representations is Reflected in Brain ResponsesRichard J. Antonello, Javier S. Turek, Vy Ai Vo, Alexander HuthNeurIPS 2021 · 60 citations
- Representation Potentials of Foundation Models for Multimodal Alignment: A SurveyJianglin Lu, Hailing Wang, Yi Xu, Yizhou Wang et al.EMNLP 2025
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
