ICLR2024

Democratizing Fine-grained Visual Recognition with Large Language Models

Mingxuan Liu, Subhankar Roy, Wenjing Li, Zhun Zhong, Nicu Sebe, Elisa Ricci

27 citations

Abstract

Well. Could you describe this photo and its wing color, head pattern, …, primary color ? B: : I see a bird in a photo. How to distinguish its specific species? Reasoning Concepts from Observations Inference … Reasoning For Each Sample : Certainly. The bird is perched on a tree branch amidst the falling snow. Its wings are grey, and it boasts a black and red pattern on its head. Notably, its dominant color is red. : Perfect. Even though I can't see it, but based on your description, I think the bird you see would be a Pyrrhuloxia, Cardinal, or Summer Tanager. Few Unlabeled Observations … VLM Test Images Semantic Classification with Reasoned Concepts Gadwall Cardinal Red Eyed Vireo Rusty Blackbird Lincoln Sparrow Tropical Kingbird … … Large Language Model Visual Question Answering Model Vision-Langauge Model Figure 1: An overview of our proposed fine-grained visual recognition (FGVR) pipeline. Left: Given few unlabelled images we exploit visual question answering (VQA) and large language models (LLM) to reason about subordinate-level category names without requiring expert knowledge. Right: At inference, we utilize the reasoned concepts to carry out FGVR via zero-shot semantic classification with a vision-language model (VLM).