MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss Alps
Valentin Gabeff, Haozhe Qi, Brendan Flaherty, Gencer Sumbul, Alexander Mathis, Devis Tuia
摘要
Monitoring wildlife is essential for ecology and ethology, especially in light of the increasing human impact on ecosystems. Camera traps have emerged as habitat-centric sensors enabling the study of wildlife populations at scale with minimal disturbance. However, the lack of annotated video datasets limits the development of powerful video understanding models needed to process the vast amount of fieldwork data collected. To advance research in wild animal behavior monitoring we present MammAlps, a multimodal and multi-view dataset of wildlife behavior monitoring from 9 camera-traps in the Swiss National Park. Mam-mAlps contains over 14 hours of video with audio, 2D seg-mentation maps and 8.5 hours of individual tracks densely labeled for species and behavior. Based on 6'135 single animal clips, we propose the first hierarchical and multimodal animal behavior recognition benchmark using audio, video and reference scene segmentation maps as inputs. Furthermore, we also propose a second ecology-oriented benchmark aiming at identifying activities, species, number of individuals and meteorological conditions from 397 multi-view and long-term ecological events, including false positive triggers. We advocate that both tasks are complementary and contribute to bridging the gap between machine learning and ecology. Code and data are available at https://github.com/eceo-epfl/MammAlps.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and IdentificationDante Francisco Wasmuht, Otto Brookes, Maximilian Schall, Pablo Palencia 等CVPR 2026 · 被引用 9 次
- PriVi: Towards a General-Purpose Video Model for Primate Behavior in the WildFelix B. Müller, Jan Frederik Meier, Timo Lüddecke, Richard Vogg 等CVPR 2026 · 被引用 2 次
- CHIRP dataset: towards long-term, individual-level, behavioral monitoring of bird populations in the wildAlex Hoi Hang Chan, Neha Singhal, Onur Kocahan, Andrea Meltzer 等CVPR 2026 · 被引用 1 次
- Explainable Forensics of Manipulated Segments in Untrimmed Long VideosYue Feng, Jingjing Li, Qijia Lu, Wei Ji 等ICML 2026
- EthoCLIP: Ontology-Enhanced Video-Language Pretraining for Animal Behavior UnderstandingYinuo Jing, Jinyan Wu, Zixi Yang, Kongming Liang 等CVPR 2026
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- Masked Autoencoders that ListenPo-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski 等NeurIPS 2022 · 被引用 524 次
- Generative Multi-View Human Action RecognitionLichen Wang, Zhengming Ding, Zhiqiang Tao, Yunyu Liu 等ICCV 2019 · 被引用 112 次
相关 Paper
- MooCap: A Multi-View Benchmark for Cow-Object-Human Interaction and Behavior DynamicsIan Noronha, Heather Neave, Upinder KaurCVPR 2026
- MammalNet: A Large-Scale Video Benchmark for Mammal Recognition and Behavior UnderstandingJun Chen, Ming Hu, Darren J. Coker, Michael L. Berumen 等CVPR 2023
- Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video UnderstandingYinuo Jing, Ruxu Zhang, Kongming Liang, Yongxiang Li 等NeurIPS 2024 · 被引用 13 次
- Animal behavioral analysis and neural encoding with transformer-based self-supervised pretrainingYanchen Wang, Han Yu, Ari Blau, Yizi Zhang 等ICLR 2026 · 被引用 8 次
- Animal Kingdom: A Large and Diverse Dataset for Animal Behavior UnderstandingXun Long Ng, Kian Eng Ong, Qichen Zheng, Yun Ni 等CVPR 2022 · 被引用 102 次
