BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment
Risa Shinoda, Kaede Shiohara, Nakamasa Inoue, Kuniaki Saito, Hiroaki Santo, Fumio Okura
Abstract
Understanding animal species from multimodal data poses an emerging challenge at the intersection of computer vision and ecology. While recent biological models, such as BioCLIP, have demonstrated strong alignment between images and textual taxonomic information for species identification, the integration of the audio modality remains an open problem. We propose BioVITA, a novel visual-textual-acoustic alignment framework for biological applications. BioVITA involves (i) a training dataset, (ii) a representation model, and (iii) a retrieval benchmark. First, we construct a large-scale training dataset comprising 1.3 million audio clips and 2.3 million images, covering 14,133 species annotated with 34 ecological trait labels. Second, building upon BioCLIP2, we introduce a two-stage training framework to effectively align audio representations with visual and textual representations. Third, we develop a cross-modal retrieval benchmark that covers all possible directional retrieval across the three modalities (i.e., image-to-audio, audio-to-text, text-to-image, and their reverse directions), with three taxonomic levels: Family, Genus, and Species. Extensive experiments demonstrate that our model learns a unified representation space that captures species-level semantics beyond taxonomy, advancing multimodal biodiversity understanding. The project page is available at: https://dahlian00.github.io/BioVITA_Page/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22043db2-e20a-42ef-a5e6-3a4537673011Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- Cross-X Learning for Fine-Grained Visual CategorizationWei Luo, Xitong Yang, Xianjie Mo, Yuheng Lu et al.ICCV 2019 · 234 citations
- Animal Kingdom: A Large and Diverse Dataset for Animal Behavior UnderstandingXun Long Ng, Kian Eng Ong, Qichen Zheng, Yun Ni et al.CVPR 2022 · 102 citations
Related papers
- BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive LearningJianyang Gu, Sam Stevens, Elizabeth G. Campolongo, Matthew J. Thompson et al.NeurIPS 2025 · 60 citations
- BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation ModelsZiheng Zhang, Xinyue Ma, Arpita Chowdhury, Elizabeth G Campolongo et al.ICLR 2026 · 3 citations
- NatureLM-audio: an Audio-Language Foundation Model for BioacousticsDavid Robinson, Marius Miron, Masato Hagiwara, Olivier PietquinICLR 2025
- BioCLIP: A Vision Foundation Model for the Tree of LifeSamuel Stevens, Jiaman Wu, Matthew J. Thompson, Elizabeth G. Campolongo et al.CVPR 2024 · 92 citations
- VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetSihan Chen, Handong Li, Qunbo Wang, Zijia Zhao et al.NeurIPS 2023 · 246 citations
