TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction
Stéphane d'Ascoli, Jérémy Rapin, Yohann Benchetrit, Hubert Banville, Jean-Remi King
摘要
Historically, neuroscience has progressed by fragmenting into specialized domains, each focusing on isolated modalities, tasks, or brain regions. While fruitful, this approach hinders the development of a unified model of cognition. Here, we introduce TRIBE, the first deep neural network trained to predict brain responses to stimuli across multiple modalities, cortical areas and individuals. By combining the pretrained representations of text, audio and video foundational models and handling their time-evolving nature with a transformer, our model can precisely model the spatial and temporal fMRI responses to videos, achieving the first place in the Algonauts 2025 brain encoding competition with a significant margin over competitors. Ablations show that while unimodal models can reliably predict their corresponding cortical networks (e.g. visual or auditory networks), they are systematically outperformed by our multimodal model in high-level associative cortices. Currently applied to perception and comprehension, our approach paves the way towards building an integrative model of representations in the human brain. Our code is available at https://anonymous.4open.science/r/algonauts-2025-C63E.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural TokensKonstantin Friedrich Willeke, Polina Turishcheva, Alex Gilbert, Goirik Chakrabarty 等ICLR 2026 · 被引用 10 次
- The Human Brain as a Dynamic Mixture of Expert Models in Video UnderstandingChristina Sartzetaki, Anne Zonneveld, Pablo Oyarzo, Alessandro T. Gifford 等ICLR 2026 · 被引用 4 次
- The Mind's Transformer: Computational Neuroanatomy of LLM-Brain AlignmentCheng-Yeh Chen, Raghupathy SivakumarICLR 2026
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- Toward a realistic model of speech processing in the brain with self-supervised learningJuliette Millet, Charlotte Caucheteux, Pierre Orhan, Yves Boubenec 等NeurIPS 2022 · 被引用 164 次
- Scaling laws for language encoding models in fMRIRichard J. Antonello, Aditya R. Vaidya, Alexander HuthNeurIPS 2023 · 被引用 137 次
相关 Paper
- Multi-modal brain encoding models for multi-modal stimuliSubba Reddy Oota, Khushbu Pahwa, Mounika Marreddy, Maneesh Kumar Singh 等ICLR 2025
- Brain encoding models based on multimodal transformers can transfer across language and visionJerry Tang, Meng Du, Vy A. Vo, Vasudev Lal 等NeurIPS 2023 · 被引用 76 次
- BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and LanguageHaitao Wu, Qirui Zhang, Zhouheng Yao, Shangquan Sun 等ICML 2026
- Revealing Vision-Language Integration in the Brain with Multimodal NetworksVighnesh Subramaniam, Colin Conwell, Christopher Wang, Gabriel Kreiman 等ICML 2024 · 被引用 19 次
- BrainFLORA: Uncovering Brain Concept Representation via Multimodal Neural EmbeddingsDongyang Li, Haoyang Qin, Mingyang Wu, Chen Wei 等ACM MM 2025 · 被引用 1 次
