Scaling Vision Transformers for Functional MRI with Flat Maps
Connor Lane, Mihir Tripathy, Leema K Murali, Ratna Grandhi, Shamus Zi Yang Sim, Sam Gijsen, Debojyoti Das, Manish Ram, Utkarsh Singh, Cesar Kadir Torrico Villanueva, YUXIANG WEI, Will Beddow
摘要
We study the problem of training self-supervised foundation models for functional MRI. Our main contributions are: (1) we introduce a new model family (CortexMAE) trained using the masked autoencoder framework on 2.1K hours of open fMRI data, and (2) we release the first open evaluation suite (Brainmarks) for fMRI foundation models. Our core innovation is simple: we adapt the Vision Transformer to fMRI by first converting each 3D fMRI volume to a 2D map using a cortical flat map projection. We directly compare flat maps to both parcellation and volume-based representations. While each has its advantages, flat maps generally perform best. We perform the first systematic scaling analysis for fMRI and observe strict power law scaling, albeit with limits. Finally, we use Brainmarks to do controlled benchmark comparisons. On subject-level trait prediction, we report a challenging null result: no single model achieves clear state-of-the-art performance. Moreover, all models struggle to outperform a simple functional connectivity baseline. On cognitive state decoding, we observe more robust performance, and in this setting our CortexMAE family outperforms prior models by a large margin. Code, models, and datasets are available at https://github.com/MedARC-AI/CortexMAE and https://github.com/MedARC-AI/Brainmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Can Natural Image Autoencoders Compactly Tokenize fMRI Volumes for Long-Range Dynamics Modeling?Peter Yongho Kim, Juhyeon Park, Jungwoo Park, Jubin Choi 等CVPR 2026 · 被引用 1 次
- BrainLM: A foundation model for brain activity recordingsJosue Ortega Caro, Antonio Henrique de Oliveira Fonseca, Syed Asad Rizvi, Matteo Rosati 等ICLR 2024 · 被引用 109 次
- Are EEG Foundation Models Worth It? Comparative Evaluation with Traditional Decoders in Diverse BCI TasksLiuyin Yang, Qiang Sun, Ang Li, Marc M. Van HulleICLR 2026
- Omni-fMRI: A Universal Atlas-Free fMRI Foundation ModelMo Wang, Wenhao Ye, Junfeng Xia, Junxiang Zhang 等ICML 2026 · 被引用 7 次
- MnemoDyn: Learning Resting State Dynamics from K FMRI sequencesSourav Pal, Viet Luong, Hoseok Lee, Tingting Dan 等ICLR 2026
