BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning
Jianyang Gu, Sam Stevens, Elizabeth G. Campolongo, Matthew J. Thompson, Net Zhang, Jiaman Wu, Andrei Kopanev, Zheda Mai, Alexander E. White, James P. Balhoff, Wasila M. Dahdul, Daniel I. Rubenstein
摘要
Foundation models trained at scale exhibit remarkable emergent behaviors, learning new capabilities beyond their initial training objectives. We find such emergent behaviors in biological vision models via large-scale contrastive vision-language training. To achieve this, we first curate TreeOfLife-200M, comprising 214 million images of living organisms, the largest and most diverse biological organism image dataset to date. We then train BioCLIP 2 on TreeOfLife-200M to distinguish different species. Despite the narrow training objective, BioCLIP 2 yields extraordinary accuracy when applied to various biological visual tasks such as habitat classification and trait prediction. We identify emergent properties in the learned embedding space of BioCLIP 2. At the inter-species level, the embedding distribution of different species aligns closely with functional and ecological meanings (e.g., beak sizes and habitats). At the intra-species level, instead of being diminished, the intra-species variations (e.g., life stages and sexes) are preserved and better separated in subspaces orthogonal to inter-species distinctions. We provide formal proof and analyses to explain why hierarchical supervision and contrastive objectives encourage these emergent properties. Crucially, our results reveal that these properties become increasingly significant with larger-scale training data, leading to a biologically meaningful embedding space.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Revisiting Semi-Supervised Learning in the Era of Foundation ModelsPing Zhang, Zheda Mai, Quang-Huy Nguyen, Wei-Lun ChaoNeurIPS 2025 · 被引用 9 次
- AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation ModelsZheda Mai, Arpita Chowdhury, Zihe Wang, Sooyoung Jeon 等CVPR 2026 · 被引用 7 次
- The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual RecognitionYuwen Tan, Yuan Qing, Boqing GongCVPR 2026 · 被引用 6 次
- Modality Alignment across Trees on Heterogeneous Hyperbolic ManifoldsWei Wu, Xiaomeng Fan, Yuwei Wu, Zhi Gao 等ICLR 2026 · 被引用 3 次
- BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation ModelsZiheng Zhang, Xinyue Ma, Arpita Chowdhury, Elizabeth G Campolongo 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
相关 Paper
- BioCLIP: A Vision Foundation Model for the Tree of LifeSamuel Stevens, Jiaman Wu, Matthew J. Thompson, Elizabeth G. Campolongo 等CVPR 2024 · 被引用 92 次
- BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic AlignmentRisa Shinoda, Kaede Shiohara, Nakamasa Inoue, Kuniaki Saito 等CVPR 2026
- DALIP: Distribution Alignment-Based Language-Image Pre-Training for Domain-Specific DataJunjie Wu, Jiangtao Xie, Zhaolin Zhang, Qilong Wang 等ICCV 2025 · 被引用 2 次
- Reproducible Scaling Laws for Contrastive Language-Image LearningMehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman 等CVPR 2023
- EthoCLIP: Ontology-Enhanced Video-Language Pretraining for Animal Behavior UnderstandingYinuo Jing, Jinyan Wu, Zixi Yang, Kongming Liang 等CVPR 2026
