Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
Yunqi Hong, Sohyun An, Andrew Bai, Neil Y. C. Lin, Cho-Jui Hsieh
Abstract
Despite Multimodal Large Language Models (MLLMs) showing promising results on general zero-shot image classification tasks, fine-grained image classification remains challenging. It demands precise attention to subtle visual details to distinguish between visually similar subcategories-details that MLLMs may easily overlook without explicit guidance. To address this, we introduce AutoSEP, an iterative self-supervised prompt learning framework designed to enhance MLLM fine-grained classification capabilities in a fully unsupervised manner. Our core idea is to leverage unlabeled data to learn a description prompt that guides MLLMs in identifying crucial discriminative features within an image, and boosts classification accuracy. We developed an automatic self-enhancing prompt learning framework called AutoSEP to iteratively improve the description prompt using unlabeled data, based on instance-level classification scoring function. AutoSEP only requires black-box access to MLLMs, eliminating the need for any training or fine-tuning. We evaluate our approach on multiple fine-grained classification datasets. It consistently outperforms other unsupervised baselines, demonstrating the effectiveness of our self-supervised optimization framework. Notably, AutoSEP in average improves 13% over standard zero-shot classification and 3% over the best-performing baselines. Code is available at https://github.com/yq-hong/AutoSEP. Image MLLM Classification Prediction Zero-shot AutoSEP Image MLLM Description Generation Description MLLM Classification Prediction Optimized Description Generation Prompt Analyze the bird depicted in the image, focusing on the following details: * Beak: Describe its color, length, shape (e.g., curved, pointed, blunt), and any distinctive markings. * Head: Describe its color, any distinctive markings (e.g., eye stripe, eyebrow, throat patch), and the presence or absence of a crest. * Wings: Note the color, pattern (e.g., banded, barred, spotted), and any white patches or bands. * Tail: Describe its length, shape (e.g., graduated, rounded, square), and any distinctive patterns (e.g., banding, barring). * Body: Describe the overall body shape (e.g., slender, stout), the color of the chest and belly, and any distinctive markings. * Overall Size and Posture: Mention if the bird appears large or small, and describe its posture (e.g., upright, hunched, relaxed). Exclude any background details or other information from your description. Optimizing with unlabeled data
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f52d09e5-c410-47c6-bcf7-953ef26034a4Cited by top-tier papers1
Ask how each one uses itBuilds on21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- ProAPO: Progressively Automatic Prompt Optimization for Visual ClassificationXiangyan Qu, Gaopeng Gou, Jiamin Zhuang, Jing Yu et al.CVPR 2025
- LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsMuhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Horst Possegger et al.NeurIPS 2023 · 63 citations
- Improved Zero-Shot Classification by Adapting VLMs with Text DescriptionsOindrila Saha, Grant Van Horn, Subhransu MajiCVPR 2024 · 26 citations
- Language-driven Fine-grained RetrievalShijie Wang, Xin Yu, Yadan Luo, Zijian Wang et al.CVPR 2026 · 2 citations
- Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought ReasoningHulingxiao He, Zijun Geng, Yuxin PengICLR 2026 · 12 citations
