Benchmarking In-the-Wild Multimodal Disease Recognition and A Versatile Baseline
Tianqi Wei, Zhi Chen, Zi Huang, Xin Yu
Abstract
Existing plant disease classification models have achieved remarkable performance in recognizing in-laboratory diseased images. However, their performance often significantly degrades in classifying in-the-wild images. Furthermore, we observed that in-the-wild plant images may exhibit similar appearances across various diseases (i.e., small inter-class discrepancy) while the same diseases may look quite different (i.e., large intra-class variance). Motivated by this observation, we propose an in-the-wild multimodal plant disease recognition dataset that contains the largest number of disease classes but also text-based descriptions for each disease. Particularly, the newly provided text descriptions are introduced to provide rich information in textual modality and facilitate in-thewild disease classification with small inter-class discrepancy and large intra-class variance issues. Therefore, our proposed dataset can be regarded as an ideal testbed for evaluating disease recognition methods in the real world. In addition, we further present a strong yet versatile baseline that models text descriptions and visual data through multiple prototypes for a given class. By fusing the contributions of multimodal prototypes in classification, our baseline can effectively address the small inter-class discrepancy and large intra-class variance issues. Remarkably, our baseline model can not only classify diseases but also recognize diseases in few-shot or training-free scenarios. Extensive benchmarking results demonstrate that our proposed in-the-wild multimodal dataset sets many new challenges to the plant disease recognition task and there is a large space to improve for future works. Our work is available at https://tqwei05.github.io/PlantWild.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3b560da-1c12-444d-8cba-a0bea6274dcbCited by top-tier papers2
- AgroBench: Vision-Language Model Benchmark in AgricultureRisa Shinoda, Nakamasa Inoue, Hirokatsu Kataoka, Masaki Onishi et al.ICCV 2025 · 10 citations
- SVIP: Semantically Contextualized Visual Patches for Zero-Shot LearningZhi Chen, Zecheng Zhao, Jingcai Guo, Jingjing Li et al.ICCV 2025 · 8 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Plant Taxonomy Meets Plant Counting: A Fine-Grained, Taxonomic Dataset for Counting Hundreds of Plant SpeciesJinyu Xu, Tianqi Hu, Xiaonan Hu, Letian Zhou et al.CVPR 2026 · 2 citations
- Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few ExamplesEleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin et al.ICLR 2020 · 692 citations
- MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from TextbooksWenqi Zeng, Yuqi Sun, Chenxi Ma, Weimin Tan et al.ACM MM 2025 · 5 citations
- Animal Kingdom: A Large and Diverse Dataset for Animal Behavior UnderstandingXun Long Ng, Kian Eng Ong, Qichen Zheng, Yun Ni et al.CVPR 2022 · 102 citations
- IRGPT: Understanding Real-World Infrared Image with Bi-Cross-Modal Curriculum on Large-Scale BenchmarkZhe Cao, Jin Zhang, Ruiheng ZhangICCV 2025 · 2 citations
