EthoCLIP: Ontology-Enhanced Video-Language Pretraining for Animal Behavior Understanding
Yinuo Jing, Jinyan Wu, Zixi Yang, Kongming Liang, Xiatian Zhu, Zhanyu Ma
Abstract
Vision-language models (VLMs) have achieved remarkable success across numerous domains, yet they lag significantly in animal behavior understanding due to severe data scarcity. Annotated animal behavior videos are prohibitively expensive and time-consuming to collect, requiring domain expertise and controlled observation conditions. To address this challenge, we leverage structured domain knowledge as an inductive bias from the Neuro Behavior Ontology (NBO), which provides professional annotations, hierarchical behavior structures, and comprehensive semantic coverage. We construct Animal-Band, an NBO-consistent dataset integrating 74,671 videos across multiple species and behaviors with semantic standardization and extended knowledge. Based on this resource, we present EthoCLIP, an ontology-enhanced vision-language contrastive learning framework that embeds ontology semantics through an ontology-aware graph module to capture hierarchical relationships among behaviors and learn structured semantic dependencies. Incorporating ontological information reduces reliance on purely datadriven learning, thereby alleviating needs for large-scale datasets. Extensive experiments validate both our dataset and method. Results demonstrate that EthoCLIP pretrained on AnimalBand substantially improves behavior recognition accuracy and transfer learning performance across diverse benchmarks, confirming that ontology-driven semantic enrichment effectively mitigates data scarcity in animal behavior understanding. Our data and code will be released at https://github.com/PRIS-CV /Ani malBand.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af3232b5-bf59-4d09-85d2-52798fdac3a8Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- Animal Kingdom: A Large and Diverse Dataset for Animal Behavior UnderstandingXun Long Ng, Kian Eng Ong, Qichen Zheng, Yun Ni et al.CVPR 2022 · 102 citations
- UniFormerV2: Unlocking the Potential of Image ViTs for Video UnderstandingKunchang Li, Yali Wang, Yinan He, Yizhuo Li et al.ICCV 2023 · 85 citations
Related papers
- Category-Specific Prompts for Animal Action Recognition with Pretrained Vision-Language ModelsYinuo Jing, Chunyu Wang, Ruxu Zhang, Kongming Liang et al.ACM MM 2023 · 6 citations
- Logic Unseen: Revealing the Logical Blindspots of Vision-Language ModelsYuchen Zhou, Jiayu Tang, Shuo Yang, Xiaoyan Xiao et al.AAAI 2026 · 2 citations
- BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive LearningJianyang Gu, Sam Stevens, Elizabeth G. Campolongo, Matthew J. Thompson et al.NeurIPS 2025 · 60 citations
- Animal behavioral analysis and neural encoding with transformer-based self-supervised pretrainingYanchen Wang, Han Yu, Ari Blau, Yizi Zhang et al.ICLR 2026 · 8 citations
- Verbs in Action: Improving verb understanding in video-language modelsLiliane Momeni, Mathilde Caron, Arsha Nagrani, Andrew Zisserman et al.ICCV 2023 · 93 citations
