The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and Identification
Dante Francisco Wasmuht, Otto Brookes, Maximilian Schall, Pablo Palencia, Christopher Beirne, Tilo Burghardt, Majid Mirmehdi, Hjalmar S. Kühl, Mimi Arandjelovic, Sam Pottie, Peter Bermant, Brandon Asheim
摘要
Automated video analysis is critical for wildlife conservation. A foundational task in this domain is multi-animal tracking (MAT), which underpins applications such as individual re-identification and behavior recognition. However, existing datasets are limited in scale, constrained to a few species, or lack sufficient temporal and geographical diversity - leaving no suitable benchmark for training general-purpose MAT models applicable across wild animal populations. To address this, we introduce SA-FARI, the largest open-source MAT dataset for wild animals. It comprises 11,609 camera trap videos collected over approximately 10 years (2014-2024) from 741 locations across 4 continents, spanning 99 species categories. Each video is exhaustively annotated culminating in 46 hours of densely annotated footage containing 16,224 masklet identities and 942,702 individual bounding boxes, segmentation masks, and species labels. Alongside the task-specific annotations, we publish anonymized camera trap locations for each video. Finally, we present comprehensive benchmarks on SA-FARI using state-of-the-art vision-language models for detection and tracking, including SAM 3, evaluated with both species-specific and generic animal prompts. We also compare against vision-only methods developed specifically for wildlife analysis. SA-FARI is the first large-scale dataset to combine high species diversity, multi-region coverage, and high-quality spatio-temporal annotations, offering a new foundation for advancing generalizable multianimal tracking in the wild. The dataset is available at https://www.conservationxlabs.com/sa-fari.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 被引用 298 次
- Animal Kingdom: A Large and Diverse Dataset for Animal Behavior UnderstandingXun Long Ng, Kian Eng Ong, Qichen Zheng, Yun Ni 等CVPR 2022 · 被引用 102 次
- BioCLIP: A Vision Foundation Model for the Tree of LifeSamuel Stevens, Jiaman Wu, Matthew J. Thompson, Elizabeth G. Campolongo 等CVPR 2024 · 被引用 92 次
- LoTE-Animal: A Long Time-span Dataset for Endangered Animal Behavior UnderstandingDan Liu, Jin Hou, Shaoli Huang, Jing Liu 等ICCV 2023 · 被引用 40 次
相关 Paper
- MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss AlpsValentin Gabeff, Haozhe Qi, Brendan Flaherty, Gencer Sumbul 等CVPR 2025
- Animal3D: A Comprehensive Dataset of 3D Animal Pose and ShapeJiacong Xu, Yi Zhang, Jiawei Peng, Wufei Ma 等ICCV 2023 · 被引用 55 次
- ATRW: A Benchmark for Amur Tiger Re-identification in the WildShuyuan Li, Jianguo Li, Hanlin Tang, Rui Qian 等ACM MM 2020 · 被引用 116 次
- Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video UnderstandingYinuo Jing, Ruxu Zhang, Kongming Liang, Yongxiang Li 等NeurIPS 2024 · 被引用 13 次
- MammalNet: A Large-Scale Video Benchmark for Mammal Recognition and Behavior UnderstandingJun Chen, Ming Hu, Darren J. Coker, Michael L. Berumen 等CVPR 2023
