Modeling the Brain’s Grammar: ROI-Guided fMRI Pretraining for Transferable and Interpretable Vision Decoding
Yulong Liu, Hua Xu, Yiyang Cai, Chunyang Jiang, Sirui Han, Yike Guo
Abstract
Recent advances in fMRI pretraining have significantly improved visual decoding accuracy by leveraging cross-subject neuroimaging datasets. A prevailing strategy aligns individual fMRI signals into a shared feature space using subject-specific adapters, followed by a shared decoder. However, this unstructured feature space overlooks the redundancy and functional correlations among voxels and fails to incorporate the brain’s intrinsic functional architecture centered on regions of interest (ROIs).To address these limitations, we propose ROITok, an ROI-guided fMRI pretraining framework. Our method introduces Sparse ROI Context Fusion to learn ROI-level visual representations and captures functional synergy between ROIs from cross-subject data. Inspired by Matryoshka Representation Learning (MRL), we design an embedding compression scheme that prioritizes the most informative visual components first, with later tokens adding progressively finer but still useful details. ROITok achieves strong transfer learning performance on the NSD and GOD datasets and shows strong resilience against high-level additive noises, while offering better interpretability and enabling new applications. It allows for quantitative assessment of each brain region’s contribution to decoding tasks. Our analysis shows that ROI-based pretraining can automatically learn the brain’s visual hierarchy. Different ROIs can provide complementary contexts for decoding tasks; combining them improves decoding robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ce79986-420a-4535-8ea7-a26428496bacBuilds on13
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- UniPC: A Unified Predictor-Corrector Framework for Fast Sampling of Diffusion ModelsWenliang Zhao, Lujia Bai, Yongming Rao, Jie Zhou et al.NeurIPS 2023 · 537 citations
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford et al.NeurIPS 2022 · 364 citations
- Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion PriorsPaul S. Scotti, Atmadeep Banerjee, Jimmie Goode, Stepan Shabalin et al.NeurIPS 2023 · 282 citations
- Versatile Diffusion: Text, Images and Variations All in One Diffusion ModelXingqian Xu, Zhangyang Wang, Eric J. Zhang, Kai Wang et al.ICCV 2023 · 265 citations
Related papers
- Meta-Learning In-Context Enables Training-Free Cross Subject Brain DecodingMu Nan, Muquan Yu, Weijian Mai, Jacob S. Prince et al.CVPR 2026 · 2 citations
- Learning Brain Representation with Hierarchical Visual EmbeddingsJiawen Zheng, Haonan Jia, MING LI, Yuhui Zheng et al.ICLR 2026 · 3 citations
- Orthogonal Contrastive Learning for Multi-Representation fMRI AnalysisTony YousefnezhadNeurIPS 2025 · 1 citation
- See Through Their Minds: Learning Transferable Brain Decoding Models from Cross-Subject fMRIYulong Liu, Yongqiang Ma, Guibo Zhu, Haodong Jing et al.AAAI 2025 · 8 citations
- Meta-Learning an In-Context Transformer Model of Human Higher Visual CortexMuquan Yu, Mu Nan, Hossein Adeli, Jacob S. Prince et al.NeurIPS 2025 · 5 citations
