BioCLIP: A Vision Foundation Model for the Tree of Life
Samuel Stevens, Jiaman Wu, Matthew J. Thompson, Elizabeth G. Campolongo, Chan Hee Song, David Edward Carlyn, Li Dong, Wasila M. Dahdul, Charles V. Stewart, Tanya Y. Berger-Wolf, Wei-Lun Chao, Yu Su
摘要
Images of the natural world, collected by a variety of cameras, from drones to individual phones, are increasingly abundant sources of biological information. There is an ex-plosion of computational methods and tools, particularly computer vision, for extracting biologically relevant information from images for science and conservation. Yet most of these are bespoke approaches designed for a specific task and are not easily adaptable or extendable to new questions, contexts, and datasets. A vision model for general or-ganismal biology questions on images is of timely need. To approach this, we curate and release Tree Of Life-10m, the largest and most diverse ML-ready dataset of biology images. We then develop Bioclip, a foundation model for the tree of life, leveraging the unique properties of bi-ology captured by Treeoflife-10m, namely the abun-dance and variety of images of plants, animals, and fungi, together with the availability of rich structured biological knowledge. We rigorously benchmark our approach on di-verse fine-grained biology classification tasks and find that BloCLIP consistently and substantially outperforms existing baselines (by 16% to 17% absolute). Intrinsic evaluation reveals that BloCLIP has learned a hierarchical representation conforming to the tree of life, shedding light on its strong generalizability.<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>imageomics.github.io/bioclip has models, data and code.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper46
- BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive LearningJianyang Gu, Sam Stevens, Elizabeth G. Campolongo, Matthew J. Thompson 等NeurIPS 2025 · 被引用 60 次
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryYu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang 等EMNLP 2024 · 被引用 28 次
- CLIP-Guided Backdoor Defense through Entropy-Based Poisoned Dataset SeparationBinyan Xu, Fan Yang, Xilin Dai, Di Tang 等ACM MM 2025 · 被引用 12 次
- AVEX: What Matters for Animal Vocalization EncodingMarius Miron, David Robinson, Milad Alizadeh, Ellen Gilsenan-McMahon 等ICLR 2026 · 被引用 9 次
- The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and IdentificationDante Francisco Wasmuht, Otto Brookes, Maximilian Schall, Pablo Palencia 等CVPR 2026 · 被引用 9 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic AlignmentRisa Shinoda, Kaede Shiohara, Nakamasa Inoue, Kuniaki Saito 等CVPR 2026
- BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation ModelsZiheng Zhang, Xinyue Ma, Arpita Chowdhury, Elizabeth G Campolongo 等ICLR 2026 · 被引用 3 次
- BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific LiteratureAlejandro Lozano, Min Woo Sun, James Burgess, Liangyu Chen 等CVPR 2025
- DALIP: Distribution Alignment-Based Language-Image Pre-Training for Domain-Specific DataJunjie Wu, Jiangtao Xie, Zhaolin Zhang, Qilong Wang 等ICCV 2025 · 被引用 2 次
- Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal ModelsHulingxiao He, Zhi Tan, Yuxin PengCVPR 2026 · 被引用 3 次
