SwiFT: Swin 4D fMRI Transformer
Peter Yongho Kim, Junbeom Kwon, Sunghwan Joo, Sangyoon Bae, Donggyu Lee, Yoonho Jung, Shinjae Yoo, Jiook Cha, Taesup Moon
摘要
Modeling spatiotemporal brain dynamics from high-dimensional data, such as functional Magnetic Resonance Imaging (fMRI), is a formidable task in neuroscience. Existing approaches for fMRI analysis utilize hand-crafted features, but the process of feature extraction risks losing essential information in fMRI scans. To address this challenge, we present SwiFT (Swin 4D fMRI Transformer), a Swin Transformer architecture that can learn brain dynamics directly from fMRI volumes in a memory and computation-efficient manner. SwiFT achieves this by implementing a 4D window multi-head self-attention mechanism and absolute positional embeddings. We evaluate SwiFT using multiple large-scale resting-state fMRI datasets, including the Human Connectome Project (HCP), Adolescent Brain Cognitive Development (ABCD), and UK Biobank (UKB) datasets, to predict sex, age, and cognitive intelligence. Our experimental outcomes reveal that SwiFT consistently outperforms recent state-of-the-art models. Furthermore, by leveraging its end-to-end learning capability, we show that contrastive loss-based self-supervised pre-training of SwiFT can enhance performance on downstream tasks. Additionally, we employ an explainable AI method to identify the brain regions associated with sex classification. To our knowledge, SwiFT is the first Swin Transformer architecture to process dimensional spatiotemporal brain functional data in an end-to-end fashion. Our work holds substantial potential in facilitating scalable learning of functional brain imaging in neuroscience research by reducing the hurdles associated with applying Transformer models to high-dimensional fMRI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Brain-JEPA: Brain Dynamics Foundation Model with Gradient Positioning and Spatiotemporal MaskingZijian Dong, Ruilin Li, Yilei Wu, Thuan Tinh Nguyen 等NeurIPS 2024 · 被引用 96 次
- fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI UnderstandingYuxiang Wei, Yanteng Zhang, Xi Xiao, Chengxuan Qian 等CVPR 2026 · 被引用 11 次
- GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence SimulationsFabian Paischer, Gianluca Galletti, William Hornsby, Paul Setinek 等NeurIPS 2025 · 被引用 9 次
- DCA: Graph-Guided Deep Embedding Clustering for Brain AtlasesMo Wang, Kaining Peng, Jingsheng Tang, Hongkai Wen 等NeurIPS 2025 · 被引用 6 次
- Can Natural Image Autoencoders Compactly Tokenize fMRI Volumes for Long-Range Dynamics Modeling?Peter Yongho Kim, Juhyeon Park, Jungwoo Park, Jubin Choi 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper8
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
相关 Paper
- CrosST: Cross Swin 4D Transformer for Multi-Modal Alzheimer's DetectionHao Wang, Hanxiao Li, Li XuACM MM 2025 · 被引用 1 次
- MnemoDyn: Learning Resting State Dynamics from K FMRI sequencesSourav Pal, Viet Luong, Hoseok Lee, Tingting Dan 等ICLR 2026
- BrainMoE: Cognition Joint Embedding via Mixture-of-Expert Towards Robust Brain Foundation ModelZiquan Wei, Tingting Dan, Tianlong Chen, Guorong WuNeurIPS 2025 · 被引用 2 次
- SIM: Surface-based fMRI Analysis for Inter-Subject Multimodal Decoding from Movie-Watching ExperimentsSimon Dahan, Gabriel Bénédict, Logan Zane John Williams, Yourong Guo 等ICLR 2025
- Scaling Vision Transformers for Functional MRI with Flat MapsConnor Lane, Mihir Tripathy, Leema K Murali, Ratna Grandhi 等ICML 2026 · 被引用 3 次
