TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction
Stéphane d'Ascoli, Jérémy Rapin, Yohann Benchetrit, Hubert Banville, Jean-Remi King
Abstract
Historically, neuroscience has progressed by fragmenting into specialized domains, each focusing on isolated modalities, tasks, or brain regions. While fruitful, this approach hinders the development of a unified model of cognition. Here, we introduce TRIBE, the first deep neural network trained to predict brain responses to stimuli across multiple modalities, cortical areas and individuals. By combining the pretrained representations of text, audio and video foundational models and handling their time-evolving nature with a transformer, our model can precisely model the spatial and temporal fMRI responses to videos, achieving the first place in the Algonauts 2025 brain encoding competition with a significant margin over competitors. Ablations show that while unimodal models can reliably predict their corresponding cortical networks (e.g. visual or auditory networks), they are systematically outperformed by our multimodal model in high-level associative cortices. Currently applied to perception and comprehension, our approach paves the way towards building an integrative model of representations in the human brain. Our code is available at https://anonymous.4open.science/r/algonauts-2025-C63E.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8830151f-c258-467b-9ee2-95add044ca33Cited by top-tier papers3
- OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural TokensKonstantin Friedrich Willeke, Polina Turishcheva, Alex Gilbert, Goirik Chakrabarty et al.ICLR 2026 · 10 citations
- The Human Brain as a Dynamic Mixture of Expert Models in Video UnderstandingChristina Sartzetaki, Anne Zonneveld, Pablo Oyarzo, Alessandro T. Gifford et al.ICLR 2026 · 4 citations
- The Mind's Transformer: Computational Neuroanatomy of LLM-Brain AlignmentCheng-Yeh Chen, Raghupathy SivakumarICLR 2026
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- Toward a realistic model of speech processing in the brain with self-supervised learningJuliette Millet, Charlotte Caucheteux, Pierre Orhan, Yves Boubenec et al.NeurIPS 2022 · 164 citations
- Scaling laws for language encoding models in fMRIRichard J. Antonello, Aditya R. Vaidya, Alexander HuthNeurIPS 2023 · 137 citations
Related papers
- Multi-modal brain encoding models for multi-modal stimuliSubba Reddy Oota, Khushbu Pahwa, Mounika Marreddy, Maneesh Kumar Singh et al.ICLR 2025
- Brain encoding models based on multimodal transformers can transfer across language and visionJerry Tang, Meng Du, Vy A. Vo, Vasudev Lal et al.NeurIPS 2023 · 76 citations
- BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and LanguageHaitao Wu, Qirui Zhang, Zhouheng Yao, Shangquan Sun et al.ICML 2026
- Revealing Vision-Language Integration in the Brain with Multimodal NetworksVighnesh Subramaniam, Colin Conwell, Christopher Wang, Gabriel Kreiman et al.ICML 2024 · 19 citations
- BrainFLORA: Uncovering Brain Concept Representation via Multimodal Neural EmbeddingsDongyang Li, Haoyang Qin, Mingyang Wu, Chen Wei et al.ACM MM 2025 · 1 citation
