Animal behavioral analysis and neural encoding with transformer-based self-supervised pretraining
Yanchen Wang, Han Yu, Ari Blau, Yizi Zhang, International Brain Laboratory, Liam Paninski, Cole Lincoln Hurwitz, Matthew R Whiteway
Abstract
The brain can only be fully understood through the lens of the behavior it generates--a guiding principle in modern neuroscience research that nevertheless presents significant technical challenges. Many studies capture behavior with cameras, but video analysis approaches typically rely on specialized models requiring extensive labeled data. We address this limitation with BEAST (BEhavioral Analysis via Self-supervised pretraining of Transformers), a novel and scalable framework that pretrains experiment-specific vision transformers for diverse neuro-behavior analyses. BEAST combines masked autoencoding with temporal contrastive learning to effectively leverage unlabeled video data. Through comprehensive evaluation across multiple species, we demonstrate improved performance in three critical neuro-behavioral tasks: extracting behavioral features that correlate with neural activity, and pose estimation and action segmentation in both the single- and multi-animal settings. Our method establishes a powerful and versatile backbone model that accelerates behavioral analysis in scenarios where labeled data remains scarce.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b4e1302-5586-49df-98ba-59bd3b16e555Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Learning Disentangled Behavior EmbeddingsChanghao Shi, Sivan Schwartz, Shahar Levy, Shay Achvat et al.NeurIPS 2021 · 13 citations
- MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations of BehaviorJennifer J. Sun, Markus Marks, Andrew Wesley Ulmer, Dipam Chakraborty et al.ICML 2023 · 20 citations
- EthoCLIP: Ontology-Enhanced Video-Language Pretraining for Animal Behavior UnderstandingYinuo Jing, Jinyan Wu, Zixi Yang, Kongming Liang et al.CVPR 2026
- Long-Short Temporal Contrastive Learning of Video TransformersJue Wang, Gedas Bertasius, Du Tran, Lorenzo TorresaniCVPR 2022 · 44 citations
- Category-Specific Prompts for Animal Action Recognition with Pretrained Vision-Language ModelsYinuo Jing, Chunyu Wang, Ruxu Zhang, Kongming Liang et al.ACM MM 2023 · 6 citations
