OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens
Konstantin Friedrich Willeke, Polina Turishcheva, Alex Gilbert, Goirik Chakrabarty, Hasan Atakan Bedel, Paul G. Fahey, Yongrong Qiu, Marissa A. Weis, Michaela Vystrčilová, Taliah Muhammad, Lydia Ntanavara, Rachel E Froebe
Abstract
Scaling data and artificial neural networks has transformed AI, driving breakthroughs in language and vision. Whether similar principles apply to modeling brain activity remains unclear. Here we leveraged a dataset of 3.1 million neurons from the visual cortex of 73 mice across 323 sessions, totaling more than 150 billion neural tokens recorded during natural movies, images and parametric stimuli, and behavior. We train multi-modal, multi-task models that support three regimes flexibly at test time: neural prediction, behavioral decoding, neural forecasting, or any combination of the three. OmniMouse achieves state-of-the-art performance, outperforming specialized baselines across nearly all evaluation regimes. We find that performance scales reliably with more data, but gains from increasing model size saturate. This inverts the standard AI scaling story: in language and computer vision, massive datasets make parameter scaling the primary driver of progress, whereas in brain modeling -even in the mouse visual cortex, a relatively simple system -models remain data-limited despite vast recordings. The observation of systematic scaling raises the possibility of phase transitions in neural modeling, where larger and richer datasets might unlock qualitatively new capabilities, paralleling the emergent properties seen in large language models. Code available at https://github.com/enigma-brain/omnimouse . Figure 1: A. OmniMouse unifies neural prediction, behavior decoding, and forecasting tasks. B. Scaling model size on 150+ billion neural tokens shows performance saturation, unlike language models. C. In contrast, scaling data consistently improves performance across all model sizes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9703b357-97ad-4bb0-8e3d-81fd93159028Builds on32
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- Hiera: A Hierarchical Vision Transformer without the Bells-and-WhistlesChaitanya Ryali, Yuan-Ting Hu, Daniel Bolya, Chen Wei et al.ICML 2023 · 388 citations
- BIOT: Biosignal Transformer for Cross-data Learning in the WildChaoqi Yang, M. Brandon Westover, Jimeng SunNeurIPS 2023 · 345 citations
- Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCIWei-Bang Jiang, Li-Ming Zhao, Bao-Liang LuICLR 2024 · 298 citations
- Brain Network TransformerXuan Kan, Wei Dai, Hejie Cui, Zilong Zhang et al.NeurIPS 2022 · 272 citations
Related papers
- Multi-session, multi-task neural decoding from distinct cell-types and brain regionsMehdi Azabou, Krystal Xuejing Pan, Vinam Arora, Ian Jarratt Knight et al.ICLR 2025
- Scaling laws for language encoding models in fMRIRichard J. Antonello, Aditya R. Vaidya, Alexander HuthNeurIPS 2023 · 137 citations
- Neural Regression, Representational Similarity, Model Zoology & Neural Taskonomy at Scale in Rodent Visual CortexColin Conwell, David Mayo, Andrei Barbu, Michael A. Buice et al.NeurIPS 2021 · 31 citations
- Multimodal Scaling Laws for Task & Data-Optimized Models of Visual CortexAbdülkadir Gökce, Yingtian Tang, Martin SchrimpfICML 2026
- Neuroformer: Multimodal and Multitask Generative Pretraining for Brain DataAntonis Antoniades, Yiyi Yu, Joseph Canzano, William Yang Wang et al.ICLR 2024 · 21 citations
