Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
Yunqi Hong, Sohyun An, Andrew Bai, Neil Y. C. Lin, Cho-Jui Hsieh
摘要
Despite Multimodal Large Language Models (MLLMs) showing promising results on general zero-shot image classification tasks, fine-grained image classification remains challenging. It demands precise attention to subtle visual details to distinguish between visually similar subcategories-details that MLLMs may easily overlook without explicit guidance. To address this, we introduce AutoSEP, an iterative self-supervised prompt learning framework designed to enhance MLLM fine-grained classification capabilities in a fully unsupervised manner. Our core idea is to leverage unlabeled data to learn a description prompt that guides MLLMs in identifying crucial discriminative features within an image, and boosts classification accuracy. We developed an automatic self-enhancing prompt learning framework called AutoSEP to iteratively improve the description prompt using unlabeled data, based on instance-level classification scoring function. AutoSEP only requires black-box access to MLLMs, eliminating the need for any training or fine-tuning. We evaluate our approach on multiple fine-grained classification datasets. It consistently outperforms other unsupervised baselines, demonstrating the effectiveness of our self-supervised optimization framework. Notably, AutoSEP in average improves 13% over standard zero-shot classification and 3% over the best-performing baselines. Code is available at https://github.com/yq-hong/AutoSEP. Image MLLM Classification Prediction Zero-shot AutoSEP Image MLLM Description Generation Description MLLM Classification Prediction Optimized Description Generation Prompt Analyze the bird depicted in the image, focusing on the following details: * Beak: Describe its color, length, shape (e.g., curved, pointed, blunt), and any distinctive markings. * Head: Describe its color, any distinctive markings (e.g., eye stripe, eyebrow, throat patch), and the presence or absence of a crest. * Wings: Note the color, pattern (e.g., banded, barred, spotted), and any white patches or bands. * Tail: Describe its length, shape (e.g., graduated, rounded, square), and any distinctive patterns (e.g., banding, barring). * Body: Describe the overall body shape (e.g., slender, stout), the color of the chest and belly, and any distinctive markings. * Overall Size and Posture: Mention if the bird appears large or small, and describe its posture (e.g., upright, hunched, relaxed). Exclude any background details or other information from your description. Optimizing with unlabeled data
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
相关 Paper
- ProAPO: Progressively Automatic Prompt Optimization for Visual ClassificationXiangyan Qu, Gaopeng Gou, Jiamin Zhuang, Jing Yu 等CVPR 2025
- LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsMuhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Horst Possegger 等NeurIPS 2023 · 被引用 63 次
- Improved Zero-Shot Classification by Adapting VLMs with Text DescriptionsOindrila Saha, Grant Van Horn, Subhransu MajiCVPR 2024 · 被引用 26 次
- Language-driven Fine-grained RetrievalShijie Wang, Xin Yu, Yadan Luo, Zijian Wang 等CVPR 2026 · 被引用 2 次
- Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought ReasoningHulingxiao He, Zijun Geng, Yuxin PengICLR 2026 · 被引用 12 次
