MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss Alps
Valentin Gabeff, Haozhe Qi, Brendan Flaherty, Gencer Sumbul, Alexander Mathis, Devis Tuia
Abstract
Monitoring wildlife is essential for ecology and ethology, especially in light of the increasing human impact on ecosystems. Camera traps have emerged as habitat-centric sensors enabling the study of wildlife populations at scale with minimal disturbance. However, the lack of annotated video datasets limits the development of powerful video understanding models needed to process the vast amount of fieldwork data collected. To advance research in wild animal behavior monitoring we present MammAlps, a multimodal and multi-view dataset of wildlife behavior monitoring from 9 camera-traps in the Swiss National Park. Mam-mAlps contains over 14 hours of video with audio, 2D seg-mentation maps and 8.5 hours of individual tracks densely labeled for species and behavior. Based on 6'135 single animal clips, we propose the first hierarchical and multimodal animal behavior recognition benchmark using audio, video and reference scene segmentation maps as inputs. Furthermore, we also propose a second ecology-oriented benchmark aiming at identifying activities, species, number of individuals and meteorological conditions from 397 multi-view and long-term ecological events, including false positive triggers. We advocate that both tasks are complementary and contribute to bridging the gap between machine learning and ecology. Code and data are available at https://github.com/eceo-epfl/MammAlps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 565ddddd-0aba-4576-8220-4685e343449eCited by top-tier papers5
- The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and IdentificationDante Francisco Wasmuht, Otto Brookes, Maximilian Schall, Pablo Palencia et al.CVPR 2026 · 9 citations
- PriVi: Towards a General-Purpose Video Model for Primate Behavior in the WildFelix B. Müller, Jan Frederik Meier, Timo Lüddecke, Richard Vogg et al.CVPR 2026 · 2 citations
- CHIRP dataset: towards long-term, individual-level, behavioral monitoring of bird populations in the wildAlex Hoi Hang Chan, Neha Singhal, Onur Kocahan, Andrea Meltzer et al.CVPR 2026 · 1 citation
- Explainable Forensics of Manipulated Segments in Untrimmed Long VideosYue Feng, Jingjing Li, Qijia Lu, Wei Ji et al.ICML 2026
- EthoCLIP: Ontology-Enhanced Video-Language Pretraining for Animal Behavior UnderstandingYinuo Jing, Jinyan Wu, Zixi Yang, Kongming Liang et al.CVPR 2026
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Masked Autoencoders that ListenPo-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski et al.NeurIPS 2022 · 524 citations
- Generative Multi-View Human Action RecognitionLichen Wang, Zhengming Ding, Zhiqiang Tao, Yunyu Liu et al.ICCV 2019 · 112 citations
Related papers
- MooCap: A Multi-View Benchmark for Cow-Object-Human Interaction and Behavior DynamicsIan Noronha, Heather Neave, Upinder KaurCVPR 2026
- MammalNet: A Large-Scale Video Benchmark for Mammal Recognition and Behavior UnderstandingJun Chen, Ming Hu, Darren J. Coker, Michael L. Berumen et al.CVPR 2023
- Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video UnderstandingYinuo Jing, Ruxu Zhang, Kongming Liang, Yongxiang Li et al.NeurIPS 2024 · 13 citations
- Animal behavioral analysis and neural encoding with transformer-based self-supervised pretrainingYanchen Wang, Han Yu, Ari Blau, Yizi Zhang et al.ICLR 2026 · 8 citations
- Animal Kingdom: A Large and Diverse Dataset for Animal Behavior UnderstandingXun Long Ng, Kian Eng Ong, Qichen Zheng, Yun Ni et al.CVPR 2022 · 102 citations
