CLIBD: Bridging Vision and Genomics for Biodiversity Monitoring at Scale
ZeMing Gong, Austin T. Wang, Xiaoliang Huo, Joakim Bruslund Haurum, Scott C. Lowe, Graham W. Taylor, Angel X. Chang
摘要
Measuring biodiversity is crucial for understanding ecosystem health. While prior works have developed machine learning models for taxonomic classification of photographic images and DNA separately, in this work, we introduce a multimodal approach combining both, using CLIP-style contrastive learning to align images, barcode DNA, and text-based representations of taxonomic labels in a unified embedding space. This allows for accurate classification of both known and unknown insect species without task-specific fine-tuning, leveraging contrastive learning for the first time to fuse barcode DNA and image data. Our method surpasses previous single-modality approaches in accuracy by over 8% on zero-shot learning tasks, showcasing its effectiveness in biodiversity studies. Recently, BioCLIP [61] used CLIP-style contrastive learning [50] to align images with common names and taxonomic descriptions to classify plants, animals, and fungi. While they showed that aligning image representations to text can help improve classification, taxonomic labels, which are not always available to the species level, are needed to obtain text descriptions. In this work, we study whether, by aligning to DNA barcodes (instead of text) during pretraining, we can learn improved representations of images for use in tasks relevant to biodiversity. We propose CLIBD, which uses contrastive learning to map taxonomic labels, biological images and barcode DNA to the same embedding space. By leveraging DNA barcodes, we eliminate the reliance on manual taxonomic labels (as used for BioCLIP) while still incorporating rich taxonomic information into the representation. This is advantageous since DNA barcodes can be obtained at scale more readily than taxonomic labels, which require manual inspection from a human expert [23, 24, 60] . We also investigate leveraging partial taxonomic annotations, when available,
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation ModelsZiheng Zhang, Xinyue Ma, Arpita Chowdhury, Elizabeth G Campolongo 等ICLR 2026 · 被引用 3 次
- G2PDiffusion: Cross-Species Genotype-to-Phenotype Prediction Via Evolutionary DiffusionMengdi Liu, Zhangyang Gao, Hong Chang, Stan Z. Li 等ICCV 2025 · 被引用 2 次
- Global and Local Entailment Learning for Natural World ImagerySrikumar Sastry, Aayush Dhakal, Eric Xing, Subash Khanal 等ICCV 2025
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 被引用 907 次
- DALIP: Distribution Alignment-Based Language-Image Pre-Training for Domain-Specific DataJunjie Wu, Jiangtao Xie, Zhaolin Zhang, Qilong Wang 等ICCV 2025 · 被引用 2 次
- Understanding Transferable Representation Learning and Zero-shot Transfer in CLIPZixiang Chen, Yihe Deng, Yuanzhi Li, Quanquan GuICLR 2024 · 被引用 21 次
- CellCLIP - Learning Perturbation Effects in Cell Painting via Text-Guided Contrastive LearningMingyu Lu, Ethan Weinberger, Chanwoo Kim, Su-In LeeNeurIPS 2025 · 被引用 9 次
- CLIPPO: Image-and-Language Understanding from Pixels OnlyMichael Tschannen, Basil Mustafa, Neil HoulsbyCVPR 2023
