The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and Identification
Dante Francisco Wasmuht, Otto Brookes, Maximilian Schall, Pablo Palencia, Christopher Beirne, Tilo Burghardt, Majid Mirmehdi, Hjalmar S. Kühl, Mimi Arandjelovic, Sam Pottie, Peter Bermant, Brandon Asheim
Abstract
Automated video analysis is critical for wildlife conservation. A foundational task in this domain is multi-animal tracking (MAT), which underpins applications such as individual re-identification and behavior recognition. However, existing datasets are limited in scale, constrained to a few species, or lack sufficient temporal and geographical diversity - leaving no suitable benchmark for training general-purpose MAT models applicable across wild animal populations. To address this, we introduce SA-FARI, the largest open-source MAT dataset for wild animals. It comprises 11,609 camera trap videos collected over approximately 10 years (2014-2024) from 741 locations across 4 continents, spanning 99 species categories. Each video is exhaustively annotated culminating in 46 hours of densely annotated footage containing 16,224 masklet identities and 942,702 individual bounding boxes, segmentation masks, and species labels. Alongside the task-specific annotations, we publish anonymized camera trap locations for each video. Finally, we present comprehensive benchmarks on SA-FARI using state-of-the-art vision-language models for detection and tracking, including SAM 3, evaluated with both species-specific and generic animal prompts. We also compare against vision-only methods developed specifically for wildlife analysis. SA-FARI is the first large-scale dataset to combine high species diversity, multi-region coverage, and high-quality spatio-temporal annotations, offering a new foundation for advancing generalizable multianimal tracking in the wild. The dataset is available at https://www.conservationxlabs.com/sa-fari.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf313847-1392-4182-9b27-e0a1bf331e89Cited by top-tier papers1
Ask how each one uses itBuilds on14
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
- Animal Kingdom: A Large and Diverse Dataset for Animal Behavior UnderstandingXun Long Ng, Kian Eng Ong, Qichen Zheng, Yun Ni et al.CVPR 2022 · 102 citations
- BioCLIP: A Vision Foundation Model for the Tree of LifeSamuel Stevens, Jiaman Wu, Matthew J. Thompson, Elizabeth G. Campolongo et al.CVPR 2024 · 92 citations
- LoTE-Animal: A Long Time-span Dataset for Endangered Animal Behavior UnderstandingDan Liu, Jin Hou, Shaoli Huang, Jing Liu et al.ICCV 2023 · 40 citations
Related papers
- MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss AlpsValentin Gabeff, Haozhe Qi, Brendan Flaherty, Gencer Sumbul et al.CVPR 2025
- Animal3D: A Comprehensive Dataset of 3D Animal Pose and ShapeJiacong Xu, Yi Zhang, Jiawei Peng, Wufei Ma et al.ICCV 2023 · 55 citations
- ATRW: A Benchmark for Amur Tiger Re-identification in the WildShuyuan Li, Jianguo Li, Hanlin Tang, Rui Qian et al.ACM MM 2020 · 116 citations
- Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video UnderstandingYinuo Jing, Ruxu Zhang, Kongming Liang, Yongxiang Li et al.NeurIPS 2024 · 13 citations
- MammalNet: A Large-Scale Video Benchmark for Mammal Recognition and Behavior UnderstandingJun Chen, Ming Hu, Darren J. Coker, Michael L. Berumen et al.CVPR 2023
