Lite-Mind: Towards Efficient and Robust Brain Representation Learning
Zixuan Gong, Qi Zhang, Guangyin Bao, Lei Zhu, Yu Zhang, Ke Liu, Liang Hu, Duoqian Miao
Abstract
The limited data availability and the low signal-to-noise ratio of fMRI signals lead to the challenging task of fMRI-to-image retrieval. State-of-the-art MindEye remarkably improves fMRI-to-image retrieval performance by leveraging a large model, i.e., a 996M MLP Backbone per subject, to align fMRI embeddings to the final hidden layer of CLIP's Vision Transformer (ViT). However, significant individual variations exist among subjects, even under identical experimental setups, mandating the training of large subject-specific models. The substantial parameters pose significant challenges in deploying fMRI decoding on practical devices. To this end, we propose Lite-Mind, a lightweight, efficient, and robust brain representation learning paradigm based on Discrete Fourier Transform (DFT), which efficiently aligns fMRI voxels to fine-grained information of CLIP. We elaborately design a DFT backbone with Spectrum Compression and Frequency Projector modules to learn informative and robust voxel embeddings. Our experiments demonstrate that Lite-Mind achieves an impressive 94.6% fMRI-to-image retrieval accuracy on the NSD dataset for Subject 1, with 98.7% fewer parameters than MindEye. Lite-Mind is also proven to be able to be migrated to smaller fMRI datasets and establishes a new state-of-the-art for zero-shot classification on the GOD dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video ReconstructionZixuan Gong, Guangyin Bao, Qi Zhang, Zhongwei Wan et al.NeurIPS 2024 · 39 citations
- Enhancing Text-to-Image Diffusion Transformer via Split-Text ConditioningYu Zhang, Jialei Zhou, Xinchen Li, Qi Zhang et al.NeurIPS 2025 · 11 citations
- Wills Aligner: Multi-Subject Collaborative Brain Visual DecodingGuangyin Bao, Qi Zhang, Zixuan Gong, Jialei Zhou et al.AAAI 2025 · 10 citations
- MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic CorrectionZixuan Gong, Qi Zhang, Guangyin Bao, Lei Zhu et al.AAAI 2025 · 2 citations
- Neurons: Emulating the Human Visual Cortex Improves Fidelity and Interpretability in fMRI-to-Video ReconstructionHaonan Wang, Qixiang Zhang, Lehan Wang, Xuanqi Huang et al.ICCV 2025 · 2 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Spectral Temporal Graph Neural Network for Multivariate Time-series ForecastingDefu Cao, Yujing Wang, Juanyong Duan, Ce Zhang et al.NeurIPS 2020 · 841 citations
- Global Filter Networks for Image ClassificationYongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu et al.NeurIPS 2021 · 798 citations
- Frequency-domain MLPs are More Effective Learners in Time Series ForecastingKun Yi, Qi Zhang, Wei Fan, Shoujin Wang et al.NeurIPS 2023 · 567 citations
Related papers
- Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion PriorsPaul S. Scotti, Atmadeep Banerjee, Jimmie Goode, Stepan Shabalin et al.NeurIPS 2023 · 282 citations
- MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of DataPaul S. Scotti, Mihir Tripathy, Cesar Torrico, Reese Kneeland et al.ICML 2024 · 117 citations
- EEGiT: Teaching Vision Transformers to Understand the EEG signalJiahao Zhou, Chenghao Xu, Wei Wang, Erkun Yang et al.CVPR 2026
- AmorLIP: Efficient Language-Image Pretraining via AmortizationHaotian Sun, Yitong Li, Yuchen Zhuang, Niao He et al.NeurIPS 2025 · 2 citations
- MiniViT: Compressing Vision Transformers with Weight MultiplexingJinnian Zhang, Houwen Peng, Kan Wu, Mengchen Liu et al.CVPR 2022 · 115 citations
