MammalNet: A Large-Scale Video Benchmark for Mammal Recognition and Behavior Understanding
Jun Chen, Ming Hu, Darren J. Coker, Michael L. Berumen, Blair R. Costelloe, Sara Beery, Anna Rohrbach, Mohamed Elhoseiny
Abstract
ized annotations and therefore do not facilitate localization of targeted behaviors within longer video sequences. Thus, we propose MammalNet, a new large-scale animal behavior dataset with taxonomy-guided annotations of mammals and their common behaviors. MammalNet contains over 18K videos totaling 539 hours, which is ∼10 times larger than the largest existing animal behavior dataset [36]. It covers 17 orders, 69 families, and 173 mammal categories for animal categorization and captures 12 high-level animal behaviors that received focus in previous animal behavior studies. We establish three benchmarks on MammalNet: standard animal and behavior recognition, compositional low-shot animal and behavior recognition, and behavior detection. Our dataset and code have been made available at: https://mammal-net.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Smoke and Mirrors in Causal Downstream TasksRiccardo Cadei, Lukas Lindorfer, Sylvia Cremer, Cordelia Schmid et al.NeurIPS 2024 · 14 citations
- Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video UnderstandingYinuo Jing, Ruxu Zhang, Kongming Liang, Yongxiang Li et al.NeurIPS 2024 · 13 citations
- The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and IdentificationDante Francisco Wasmuht, Otto Brookes, Maximilian Schall, Pablo Palencia et al.CVPR 2026 · 9 citations
- BigMaQ: A Big Macaque Motion and Animation Dataset Bridging Image and 3D Pose RepresentationsLucas Martini, Alexander Lappe, Anna Bognár, Rufin Vogels et al.ICLR 2026 · 3 citations
- MOVE: Motion-Guided Few-Shot Video Object SegmentationKaining Ying, Hengrui Hu, Henghui DingICCV 2025 · 3 citations
Builds on7
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- MViTv2: Improved Multiscale Vision Transformers for Classification and DetectionYanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam et al.CVPR 2022 · 699 citations
- Cross-Domain Adaptation for Animal Pose EstimationJinkun Cao, Hongyang Tang, Haoshu Fang, Xiaoyong Shen et al.ICCV 2019 · 209 citations
- Animal Kingdom: A Large and Diverse Dataset for Animal Behavior UnderstandingXun Long Ng, Kian Eng Ong, Qichen Zheng, Yun Ni et al.CVPR 2022 · 102 citations
- CoLA: Weakly-Supervised Temporal Action Localization With Snippet Contrastive LearningCan Zhang, Meng Cao, Dongming Yang, Jie Chen et al.CVPR 2021
Related papers
- MooCap: A Multi-View Benchmark for Cow-Object-Human Interaction and Behavior DynamicsIan Noronha, Heather Neave, Upinder KaurCVPR 2026
- MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss AlpsValentin Gabeff, Haozhe Qi, Brendan Flaherty, Gencer Sumbul et al.CVPR 2025
- Animal3D: A Comprehensive Dataset of 3D Animal Pose and ShapeJiacong Xu, Yi Zhang, Jiawei Peng, Wufei Ma et al.ICCV 2023 · 55 citations
- AnimalWeb: A Large-Scale Hierarchical Dataset of Annotated Animal FacesMuhammad Haris Khan, John McDonagh, Salman H. Khan, Muhammad Shahabuddin et al.CVPR 2020
- LoTE-Animal: A Long Time-span Dataset for Endangered Animal Behavior UnderstandingDan Liu, Jin Hou, Shaoli Huang, Jing Liu et al.ICCV 2023 · 40 citations
