Neuroformer: Multimodal and Multitask Generative Pretraining for Brain Data
Antonis Antoniades, Yiyi Yu, Joseph Canzano, William Yang Wang, Spencer L. Smith
Abstract
State-of-the-art systems neuroscience experiments yield large-scale multimodal data, and these data sets require new tools for analysis. Inspired by the success of large pretrained models in vision and language domains, we reframe the analysis of large-scale, cellular-resolution neuronal spiking data into an autoregressive spatiotemporal generation problem. Neuroformer is a multimodal, multitask generative pretrained transformer (GPT) model that is specifically designed to handle the intricacies of data in systems neuroscience. It scales linearly with feature size, can process an arbitrary number of modalities, and is adaptable to downstream tasks, such as predicting behavior. We first trained Neuroformer on simulated datasets, and found that it both accurately predicted simulated neuronal circuit activity, and also intrinsically inferred the underlying neural circuit connectivity, including direction. When pretrained to decode neural responses, the model predicted the behavior of a mouse with only few-shot fine-tuning, suggesting that the model begins learning how to do so directly from the neural representations themselves, without any explicit supervision. We used an ablation study to show that joint training on neuronal responses and behavior boosted performance, highlighting the model's ability to associate behavioral and neural representations in an unsupervised manner. These findings show that Neuroformer can analyze neural datasets and their emergent properties, informing the development of models and hypotheses associated with the brain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Latent Diffusion for Neural Spiking DataJaivardhan Kapoor, Auguste Schulz, Julius Vetter, Felix Pei et al.NeurIPS 2024 · 24 citations
- OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural TokensKonstantin Friedrich Willeke, Polina Turishcheva, Alex Gilbert, Goirik Chakrabarty et al.ICLR 2026 · 10 citations
- POCO: Scalable Neural Forecasting through Population ConditioningYu Duan, Hamza Tahir Chaudhry, Misha B. Ahrens, Christopher D. Harvey et al.NeurIPS 2025 · 8 citations
- Multi-Modal Latent Variables for Cross-Individual Primary Visual Cortex Modeling and AnalysisYu Zhu, Bo Lei, Chunfeng Song, Wanli Ouyang et al.AAAI 2025 · 5 citations
- TRACE: Contrastive learning for multi-trial time series data in neuroscienceLisa Schmors, Dominic Gonschorek, Jan Niklas Böhm, Yongrong Qiu et al.NeurIPS 2025 · 3 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
Related papers
- NetFormer: An interpretable model for recovering dynamical connectivity in neuronal population dynamicsZiyu Lu, Wuwei Zhang, Trung Le, Hao Wang et al.ICLR 2025
- Multi-session, multi-task neural decoding from distinct cell-types and brain regionsMehdi Azabou, Krystal Xuejing Pan, Vinam Arora, Ian Jarratt Knight et al.ICLR 2025
- BrainBERT: Self-supervised representation learning for intracranial recordingsChristopher Wang, Vighnesh Subramaniam, Adam Uri Yaari, Gabriel Kreiman et al.ICLR 2023 · 13 citations
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie et al.ICCV 2021 · 392 citations
- Are EEG Foundation Models Worth It? Comparative Evaluation with Traditional Decoders in Diverse BCI TasksLiuyin Yang, Qiang Sun, Ang Li, Marc M. Van HulleICLR 2026
