Modeling the Brain’s Grammar: ROI-Guided fMRI Pretraining for Transferable and Interpretable Vision Decoding
Yulong Liu, Hua Xu, Yiyang Cai, Chunyang Jiang, Sirui Han, Yike Guo
摘要
Recent advances in fMRI pretraining have significantly improved visual decoding accuracy by leveraging cross-subject neuroimaging datasets. A prevailing strategy aligns individual fMRI signals into a shared feature space using subject-specific adapters, followed by a shared decoder. However, this unstructured feature space overlooks the redundancy and functional correlations among voxels and fails to incorporate the brain’s intrinsic functional architecture centered on regions of interest (ROIs).To address these limitations, we propose ROITok, an ROI-guided fMRI pretraining framework. Our method introduces Sparse ROI Context Fusion to learn ROI-level visual representations and captures functional synergy between ROIs from cross-subject data. Inspired by Matryoshka Representation Learning (MRL), we design an embedding compression scheme that prioritizes the most informative visual components first, with later tokens adding progressively finer but still useful details. ROITok achieves strong transfer learning performance on the NSD and GOD datasets and shows strong resilience against high-level additive noises, while offering better interpretability and enabling new applications. It allows for quantitative assessment of each brain region’s contribution to decoding tasks. Our analysis shows that ROI-based pretraining can automatically learn the brain’s visual hierarchy. Different ROIs can provide complementary contexts for decoding tasks; combining them improves decoding robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- UniPC: A Unified Predictor-Corrector Framework for Fast Sampling of Diffusion ModelsWenliang Zhao, Lujia Bai, Yongming Rao, Jie Zhou 等NeurIPS 2023 · 被引用 537 次
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford 等NeurIPS 2022 · 被引用 364 次
- Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion PriorsPaul S. Scotti, Atmadeep Banerjee, Jimmie Goode, Stepan Shabalin 等NeurIPS 2023 · 被引用 282 次
- Versatile Diffusion: Text, Images and Variations All in One Diffusion ModelXingqian Xu, Zhangyang Wang, Eric J. Zhang, Kai Wang 等ICCV 2023 · 被引用 265 次
相关 Paper
- Meta-Learning In-Context Enables Training-Free Cross Subject Brain DecodingMu Nan, Muquan Yu, Weijian Mai, Jacob S. Prince 等CVPR 2026 · 被引用 2 次
- Learning Brain Representation with Hierarchical Visual EmbeddingsJiawen Zheng, Haonan Jia, MING LI, Yuhui Zheng 等ICLR 2026 · 被引用 3 次
- Orthogonal Contrastive Learning for Multi-Representation fMRI AnalysisTony YousefnezhadNeurIPS 2025 · 被引用 1 次
- See Through Their Minds: Learning Transferable Brain Decoding Models from Cross-Subject fMRIYulong Liu, Yongqiang Ma, Guibo Zhu, Haodong Jing 等AAAI 2025 · 被引用 8 次
- Meta-Learning an In-Context Transformer Model of Human Higher Visual CortexMuquan Yu, Mu Nan, Hossein Adeli, Jacob S. Prince 等NeurIPS 2025 · 被引用 5 次
